Marine organism detection model training method and marine organism detection method
By using the combination of backbone network, neck network and head network in the marine biological detection model, feature extraction and fusion processing is performed, the problem of low detection accuracy caused by the complexity of the marine environment is solved, and high-accurate marine biological detection is achieved.
Patent Information
- Application Number
- CN202510076009.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, the marine environment is complex and changeable, and there are many factors that affect target detection, resulting in low detection accuracy of marine biological detection models.
A marine biological detection model training method is adopted to extract and fusion the feature of convolutional layer, convolutional attention layer and spatial pooling layer through the combination of backbone network, neck network and head network, thereby enhancing the ability to capture marine biological features.
The marine biological detection model trained through this method can improve detection accuracy, enhance the ability to capture marine biological characteristics, and is suitable for marine biological detection.
Smart Images

Figure CN120071390A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of marine biological detection, and specifically relates to a method for training a marine biological detection model and a method for marine biological detection. Background Art
[0002] Marine biological detection is the most difficult in computer vision detection technology. With the development of marine exploration, object detection plays an important role in fields such as marine resource exploration, marine environmental monitoring, and underwater archaeology. However, the marine environment is complex and changeable, and there are far more factors affecting object detection than on land, which seriously affects the accuracy of object detection. Therefore, the object detection model applied to the marine environment has the technical defect of low detection accuracy. Summary of the Invention
[0003] The purpose of this application is to overcome the deficiencies in the prior art and provide a method for training a marine biological detection model and a method for marine biological detection, which can train a marine biological detection model with high detection accuracy for marine biological detection.
[0004] The first aspect of the embodiments of this application provides a method for training a marine biological detection model, which is applied to the marine biological detection model to be trained; the marine biological detection model includes a backbone network, a neck network, and a head network connected in sequence; the backbone network includes a convolutional layer, a plurality of convolutional attention layers, and a spatial pooling layer;
[0005] The method includes:
[0006] Input the marine biological training samples into the convolutional layer for feature extraction to obtain first sample features;
[0007] Input the first sample features into the cascaded plurality of convolutional attention layers for feature extraction and feature attention processing to obtain second sample features output by each convolutional attention layer;
[0008] Input the second sample features output by the last convolutional attention layer into the spatial pooling layer for dimensionality reduction processing to obtain pooled features;
[0009] Input the pooled features and a plurality of preset second sample features into the neck network for fusion processing and spatial transformation attention processing to obtain a plurality of fusion attention features;
[0010] Input the plurality of fusion attention features into the head network for prediction learning processing to obtain the trained marine biological detection model.
[0011] The second embodiment of this application discloses a method for marine biological detection, including:
[0012] Through the above-mentioned method for training a marine organism detection model, a trained marine organism detection model is obtained.
[0013] Input the marine image to be detected into the trained marine organism detection model to obtain the marine organism detection result output by the marine organism detection model.
[0014] Compared with the related technology, the present application is applied to train a marine organism detection model. The marine organism detection model to be trained includes a backbone network, a neck network, and a head network connected in sequence. The backbone network includes a convolutional layer, multiple convolutional attention layers, and a spatial pooling layer. The training process includes: inputting the marine organism training sample into the convolutional layer for feature extraction to obtain a first sample feature; inputting the first sample feature into the cascaded multiple convolutional attention layers for feature extraction and feature attention processing to obtain second sample features output by each convolutional attention layer; inputting the second sample feature output by the last convolutional attention layer into the spatial pooling layer for dimensionality reduction processing to obtain a pooled feature; inputting the pooled feature and a plurality of preset second sample features into the neck network for fusion processing and spatial transformation attention processing to obtain multiple fusion attention features; inputting the multiple fusion attention features into the head network for prediction learning processing until the number of training times of the marine organism detection model reaches a preset number threshold to obtain a trained marine organism detection model. Based on the trained marine organism detection model with the above model structure, fusion processing and spatial change attention processing can be performed according to the second sample features and pooled features in different spatial dimensions, enhancing the ability to capture the features of marine organisms. Therefore, a marine organism detection model with high detection accuracy can be trained for marine organism detection.
[0015] In order to more clearly understand the present application, the specific implementation manners of the present application will be described below in conjunction with the accompanying drawings. Brief Description of the Drawings
[0016] Figure 1 It is a flowchart of the method for training a marine organism detection model according to an embodiment of the present application.
[0017] Figure 2 It is a schematic diagram of the model structure of a marine organism detection model according to an embodiment of the present application.
[0018] Figure 3 It is a schematic diagram of the structure of the deformation attention module of a marine organism detection model according to an embodiment of the present application.
[0019] Figure 4 It is a schematic diagram of the structure of the dual attention module of a marine organism detection model according to an embodiment of the present application.
[0020] Figure 5Structural diagram of the transform attention module of the marine organism detection model according to an embodiment of the present application.
[0021] Figure 6 Flowchart of the marine organism detection method according to an embodiment of the present application. Detailed implementation manners
[0022] To make the objectives, technical solutions and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0023] It should be clear that the described embodiments are only part of the embodiments of the present application, rather than all of them. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the embodiments of the present application.
[0024] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. In the description of the present application, it should be understood that the terms "first", "second", "third", etc. are only used to distinguish similar objects, and do not have to be used to describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances. The singular forms of "a", "the" and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. The word "if" / "when" used herein can be interpreted as "when" or "when" or "in response to a determination".
[0025] In addition, in the description of the present application, unless otherwise specified, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0026] Please refer to Figure 1-2 , Figure 1 is the training method of the marine organism detection model according to an embodiment of the present application, which is applied to the marine organism detection model to be trained. Figure 2 is the schematic diagram of the model structure of the marine organism detection model according to an embodiment of the present application. The marine organism detection model includes a backbone network (Backbone), a neck network (Neck) and a head network (Head) connected in sequence; the backbone network includes a convolutional layer, a plurality of convolutional attention layers and a spatial pooling layer.
[0027] The method includes:
[0028] S1: Input the marine organism training sample into the convolutional layer for feature extraction to obtain the first sample feature.
[0029] The marine organism training sample refers to a training sample with target annotation for a marine image. The marine image can be a seabed image directly captured by an underwater shooting device or a seabed image obtained from a frame of a seabed video.
[0030] For Figure 2 example, the convolutional kernel size of convolutional layer P1 is 3, the stride is 2, and the padding is 1.
[0031] S2: Input the first sample feature into the cascaded multiple convolutional attention layers for feature extraction and feature attention processing to obtain the second sample features output by each convolutional attention layer.
[0032] In a feasible embodiment, the convolutional attention layer includes a convolutional module and a dual attention module. As Figure 2 shown, the outputs of convolutional module P2, convolutional module P3, convolutional module P4, and convolutional module P5 are respectively connected to the inputs of the corresponding dual attention modules DAMM, and the outputs of each dual attention module are respectively connected to the inputs of the convolutional modules of the next-level convolutional attention layer.
[0033] The convolutional module performs feature extraction processing on the input to obtain the third sample feature.
[0034] The dual attention module performs feature attention processing on the third sample feature of the same convolutional attention layer to obtain the second sample feature, and the dual attention module outputs the second sample feature to the convolutional module of the next-level convolutional attention layer.
[0035] In this embodiment, the second sample features output by each convolutional attention layer are all second sample features that are successively obtained through feature extraction processing by the convolutional module and feature attention processing by the dual attention module. Therefore, the second sample features output by each convolutional attention layer are second sample features of different spatial dimensions.
[0036] S3: Input the second sample feature output by the last convolutional attention layer into the spatial pooling layer for dimensionality reduction processing to obtain the pooled feature.
[0037] Please refer to Figure 3 , in a feasible embodiment, the spatial pooling layer includes a lightweight convolutional activation module (SqueezeConv), a deformable attention module (BRDA), a pooling fusion module, and an expansion convolutional module (EXpansionConv).
[0038] The lightweight convolution activation module performs convolution processing on the second sample feature output by the last convolutional attention layer to obtain a sample activation feature, and the lightweight convolution activation module outputs the sample activation feature to the deformable attention module and the pooling fusion module respectively; the process is shown in the following formula:
[0039]
[0040] where x′ is the sample activation feature; is the convolution processing of Squeeze Conv; X B is the second sample feature output by the last convolutional attention layer; B is the batch dimension; C′ is the number of hidden channels, H is the height dimension; W is the width dimension.
[0041] The deformable attention module performs offset deformation processing, feature attention processing and fusion processing on the sample activation feature to obtain a sixth sample feature, and the deformable attention module outputs the sixth sample feature to the pooling fusion module.
[0042] The pooling fusion model performs stacking fusion and pooling fusion processing on the sample activation feature and the sixth sample feature to obtain a first pooling fusion feature.
[0043] Among them, pooling fusion is to perform at least three times of pooling dimensionality reduction on the stacking fusion result of the sample activation feature and the sixth sample feature to obtain multiple pooling results, and then fuse all the pooling results with the sample activation feature to obtain the first pooling fusion feature.
[0044] The data processing processes of the deformable attention module and the pooling fusion model are shown in the following formula:
[0045]
[0046] where x″ is the result of stacking fusion of the sample activation feature and the sixth sample feature; BRDA(x′) is the sixth sample feature;
[0047] y 1 =MaxPool(x″)
[0048] where y 1 is the pooling result of the first pooling dimensionality reduction; MaxPool(·) is the max pooling processing;
[0049] y 2 =MaxPool(y 1 )
[0050] where y 2 is the pooling result of the second pooling dimensionality reduction;
[0051] y 3 = MaxPool(y 2 )
[0052] where y 3 is the pooling result of the third pooling dimensionality reduction;
[0053]
[0054] where z is the first pooled fusion feature.
[0055] The dilated convolution module performs dilated convolution processing on the first pooled fusion feature to obtain the pooled feature.
[0056]
[0057] where Y B is the pooled feature.
[0058] In this embodiment, through the lightweight convolution activation module, the deformable attention module, the pooling fusion module, and the dilated convolution module, the features of marine organisms with different spatial dimensions can be obtained.
[0059] S4: Input the pooled feature and a plurality of preset second sample features into the neck network for fusion processing and spatial transformation attention processing to obtain a plurality of fusion attention features.
[0060] In a feasible embodiment, the neck network includes a first upsampling fusion module, a first transformation attention module, a second upsampling fusion module, a second transformation attention module, a first convolution fusion module, a third variation attention module, a second convolution fusion module, and a fourth variation attention module;
[0061] The first upsampling fusion module performs upsampling processing on the pooled feature and then fuses it with the second sample feature output by the penultimate convolution attention layer to obtain a first upsampling fusion feature;
[0062] The first transformation attention module performs variation attention processing on the first upsampling fusion feature to obtain a first transformation attention feature;
[0063] The second upsampling fusion module performs upsampling processing on the first transformation attention feature and then fuses it with the second sample feature output by the third-to-last convolution attention layer to obtain a second upsampling fusion feature;
[0064] The second transformation attention module performs variation attention processing on the first transformation attention feature to obtain a second transformation attention feature;
[0065] After the first convolution fusion module performs convolution processing on the second transformed attention feature, it is fused with the first transformed attention feature to obtain a first convolution fusion feature;
[0066] The third variation attention module performs variation attention processing on the first convolution fusion feature to obtain a third transformed attention feature;
[0067] After the second convolution fusion module performs variation attention processing on the three transformed attention features, it is fused with the pooling feature to obtain a second convolution fusion feature;
[0068] The fourth variation attention module performs variation attention processing on the second convolution fusion feature to obtain a fourth transformed attention feature;
[0069] The second transformed attention feature, the third transformed attention feature, and the fourth transformed attention feature are determined as the fusion attention features.
[0070] In this embodiment, the first upsampling fusion module, the first transformed attention module, the second upsampling fusion module, the second transformed attention module, the first convolution fusion module, the third variation attention module, the second convolution fusion module, and the fourth variation attention module are used to perform fusion processing and transformed attention processing on features of different spatial dimensions, which can enhance the ability of the neck network to capture features of marine organisms in different spatial dimensions.
[0071] S5: Input the multiple fusion attention features into the head network for prediction learning processing to obtain a trained marine organism detection model.
[0072] Among them, the head network includes multiple decoupled heads, and each decoupled head corresponds to each fusion attention feature respectively, and is used to perform prediction learning processing according to the corresponding fusion attention feature. Specifically, the decoupled head includes a regression branch module and a classification branch module. A loss function is provided in the regression branch module to calculate the position offset between the predicted box of the regression branch and the detection label of the marine organism training sample for loss calculation; for the predicted box, the classification branch model obtains the value at each position on the classifier output tensor through pooling and convolution operations, indicating the probability that the predicted box belongs to each category, and finally filters out the final detection result through non-maximum suppression.
[0073] When training the marine organism detection model, it is trained according to a preset number threshold until the training times of the marine organism detection model reach the number threshold.
[0074] Compared with the related art, the present application is applied to training a marine organism detection model. The marine organism detection model to be trained includes a backbone network, a neck network, and a head network connected in sequence. The backbone network includes a convolutional layer, multiple convolutional attention layers, and a spatial pooling layer. The training process includes: inputting the marine organism training sample into the convolutional layer for feature extraction to obtain a first sample feature; inputting the first sample feature into the cascaded multiple convolutional attention layers for feature extraction and feature attention processing to obtain second sample features output by each convolutional attention layer; inputting the second sample feature output by the last convolutional attention layer into the spatial pooling layer for dimensionality reduction processing to obtain a pooled feature; inputting the pooled feature and a plurality of preset second sample features into the neck network for fusion processing and spatial transformation attention processing to obtain a plurality of fused attention features; inputting the plurality of fused attention features into the head network for prediction learning processing until the number of training times of the marine organism detection model reaches a preset number threshold to obtain a trained marine organism detection model. Based on the trained marine organism detection model with the above model structure, fusion processing and spatial change attention processing can be performed according to the second sample features and pooled features in different spatial dimensions, enhancing the ability to capture the features of marine organisms. Therefore, a marine organism detection model with high detection accuracy can be trained for marine organism detection.
[0075] Please refer to Figure 4 , in a feasible embodiment, the dual attention module includes an enhanced convolutional block, a dual adaptive attention block, and a first stacked fusion block;
[0076] The enhanced convolutional block performs convolutional enhancement processing on the third sample feature to obtain a fourth sample feature, and the process is shown in the following formula:
[0077] ecb_output = γ ECB ·ECB(BN(X A ))
[0078] ecb_output is the fourth sample feature; XA is the third sample feature; BN(·) is the normalization operation; ECB(·) is the convolutional enhancement processing operation; γ ECB is the learning parameter of the convolutional enhancement processing.
[0079] The dual adaptive attention block performs feature attention fusion processing on the third sample feature to obtain a fifth sample feature, and the process is shown in the following formula:
[0080] attention_output = γ DAA ·DAA(BN(X A ))
[0081] Among them, attention_output is the fifth sample feature; DAA(·) is the feature attention fusion operation; γ DAA is the learning parameter of feature attention fusion.
[0082] The first stacking fusion block performs stacking fusion processing on the third sample feature, the fourth sample feature, and the fifth sample feature to obtain the second sample feature; the process is shown in the following formula:
[0083] Y A = attention_output + ecb_output + X A
[0084] Among them, Y A is the second sample feature.
[0085] In this embodiment, by performing feature attention processing on the third sample feature through the enhanced convolution block, the dual adaptive attention block, and the first stacking fusion block, the second sample feature of each convolutional attention layer can be obtained.
[0086] In a feasible embodiment, the dual adaptive attention block includes: a first convolutional activation sub-module, a second convolutional activation sub-module, a decomposed attention fusion sub-module, and a feature fusion sub-module;
[0087] The first convolutional activation sub-module performs convolutional activation processing on the third sample feature to obtain a first activation feature, including:
[0088]
[0089] Among them, X 1 is the first activation feature; X A is the third sample feature; BN(·) is the normalization operation; is the convolutional operation, normalization operation, and activation operation.
[0090] The second convolutional activation sub-module performs convolutional activation processing on the first activation feature to obtain a second activation feature, including:
[0091]
[0092] Among them, X 2 is the second activation feature; is the convolutional operation, normalization operation, and activation operation.
[0093] The decomposed attention fusion sub-module performs decomposition processing and attention fusion convolution processing on the first activation feature and the second activation feature to obtain an attention fusion feature.
[0094] Please continue to refer toFigure 4 , in a feasible embodiment, the decomposed attention fusion sub-module includes a decomposition unit, a fusion unit, an attention normalization processing unit, and an attention fusion convolution unit;
[0095] The decomposition unit is used to perform decomposition processing on the first activation feature and the second activation feature to obtain a first decomposition feature and a second decomposition feature; the following formula is included:
[0096]
[0097] Among them, are the first decomposition feature and the second decomposition feature respectively, that is, Figure 4 the outputs of SA_S1 and SA_S2 in is a 1×1 convolution operation.
[0098] The fusion unit performs fusion processing on the first decomposition feature and the second decomposition feature to obtain a decomposed fusion feature; the following formula is included:
[0099]
[0100] Among them, φ mean (·) is Figure 4 the output of Avg in
[0101]
[0102] Among them, φ max (·) is Figure 4 the output of Max in
[0103]
[0104] Among them, is the decomposed fusion feature, that is, Figure 4 the output of Attention Fusion in
[0105] The attention normalization processing unit performs attention processing and normalization processing on the decomposed fusion feature to obtain an attention normalization feature; the following formula is included:
[0106]
[0107] Among them, Q is the attention normalization feature, that is, Figure 4 the output of the Sigmoid function in
[0108] The attention fusion convolution unit performs fusion convolution processing on the first decomposition feature, the second decomposition feature, and the attention normalization feature to obtain the attention fusion feature; the following formula is included:
[0109]
[0110] Among them, sC is the attention fusion feature; is the first decomposition feature, is the second decomposition feature, and Q[:, i::] is the i-th spatial embedding weight of the corresponding channel.
[0111] The feature fusion sub-module performs fusion processing on the first activation feature, the second activation feature, and the attention fusion feature to obtain the fifth sample feature;
[0112]
[0113] Among them, attention_output is the fifth sample feature; γ DAA is the learning parameter of feature attention fusion; sC is the attention fusion feature; X 1 is the first activation feature; X 2 is the second activation feature.
[0114] Please continue to refer to Figure 3 , in a feasible embodiment, the deformable attention module includes a first attention parameter acquisition sub-module, a deformation sub-module, a second attention parameter acquisition sub-module, and a deformable attention fusion sub-module;
[0115] The first attention parameter acquisition sub-module extracts attention parameters from the sample activation feature to obtain the first attention parameter; the following formula is included:
[0116]
[0117] Among them, W q is the weight matrix, and q is the first attention parameter.
[0118] The deformation sub-module performs offset deformation processing on the sample activation feature according to the first attention parameter to obtain the offset deformation feature; the following formula is included:
[0119]
[0120] Among them, Offset is to generate two-dimensional offsets using 1x1 convolution.
[0121]
[0122] Among them, is the offset deformation feature; Γ(·) is the bilinear interpolation method; θ(r + Offset) is the clamping operation; r is the reference point grid, which can be generated by numpy.meshgrid().
[0123] The second attention parameter acquisition sub-module performs attention parameter extraction on the offset deformation feature and the first attention parameter to obtain a second attention parameter and a third attention parameter; the following formula is included:
[0124]
[0125] where is the candidate second attention parameter; is the candidate third attention parameter; W k and W v are both
[0126] Specifically, the routing calculation TopkRouting(·) can also be used to deduce the correlation degree between regions to obtain the attention index matrix, including the following formula:
[0127]
[0128] where I is the attention index matrix; TopkRouting(·) is the TopkRouting function, which is a function used to select the most important nodes or edges in a graph neural network. This function calculates the features of each node or edge and then sorts them according to the importance of these features.
[0129] Then for each candidate region, only several of the most relevant connections are retained for fine-grained tokenization, and the key-value pairs k and v are selectively collected through the method KVGather(), including the following formula:
[0130]
[0131] where is the second attention parameter; is the third attention parameter; KVGather(k, v, I) is to collect the first several key-value pairs k and v from the attention index matrix.
[0132] The deformation attention fusion sub-module performs fusion processing on the first attention parameter, the second attention parameter, and the third attention parameter to obtain the sixth sample feature, including the following formula:
[0133]
[0134] Among them, BRDA(x′) is the sixth sample feature, and Softmax(.) is used to convert a real number vector into a probability distribution.
[0135] Please refer to Figure 5 , which is the structural diagram of the transformation attention module of this application. That is, the first transformation attention module, the second transformation attention module, the third transformation attention module, and the fourth transformation attention module all adopt the structure of the transformation attention module, that is, the module structures of the first transformation attention module, the second transformation attention module, the third transformation attention module, and the fourth transformation attention module are the same. In a feasible embodiment, taking the first transformation attention module as an example, the first transformation attention module includes a channel grouping unit, a first channel transformation unit, a feature block pooling and fusion unit, a normalization and fusion unit, and a second channel transformation unit;
[0136] The channel grouping unit groups the first upsampling and fusion features according to the channel dimension to obtain a plurality of first feature groups;
[0137]
[0138] Among them, X is a plurality of first feature groups; g is the number of groups; X i is the i-th first feature group.
[0139] The first channel transformation unit sequentially transforms the plurality of first feature groups to obtain a plurality of second feature groups.
[0140] X′ = shuffle(X) = [X′ 1 ,..., X′ g
[0141] Among them, X′ is a plurality of second feature groups.
[0142] The feature block pooling and fusion unit performs block processing, pooling processing, and fusion processing on the plurality of second feature groups to obtain a second pooling and fusion feature;
[0143] The block processing is implemented by the following formula:
[0144] X′ i1 , X′ i2 = chunk(X′ i ), i = 1,..., g
[0145] Among them, X′ i1 is the first block feature; X′ i2 is the second block feature; Chunk(·) is the block processing.
[0146] The pooling processing is implemented by the following formula:
[0147] X″ i1 = W a ·P gap (X′ i1 ) + b a
[0148] Wherein, X″ i1 is the first pooling result of the first block feature; P gap (·) is the global average pooling operation, W a and b a are learning parameters;
[0149] X″ i2 = W m ·P gmp (X′ i2 ) + b m
[0150] Wherein, X″ i2 is the second pooling result of the second block feature; P gmp (·) is the global max pooling operation, W m and b m are learning parameters.
[0151] The fusion process is implemented by the following formula:
[0152] X″ Ci = Concat[X″ i1 ; X″ i2
[0153] Wherein, X″ Ci is the transformed fusion result of the first pooling result and the second pooling result.
[0154] After the normalization fusion unit normalizes the second pooling fusion feature, it fuses with the multiple first feature groups to obtain the normalized fusion feature, and the process is shown in the following formula:
[0155] X′ Ci = σ(X″ Ci )
[0156] ·X i
[0157] Wherein, X″ Ci is the normalized fusion feature; σ(·) is the linear normalization operation; X i is the i-th first feature group.
[0158] The second channel transformation unit sequentially transforms the normalized fusion feature to obtain the first transformed attention feature with the restored feature order, and the process is shown in the following formula:
[0159] X Ci = shuffle(X' Ci )
[0160] where X Ci is the first transformed attention feature.
[0161] Please refer to Figure 6 , the second embodiment of the present application discloses a marine organism detection method, including:
[0162] S6: Through the above-mentioned marine organism detection model training method, obtain the trained marine organism detection model.
[0163] S7: Input the marine image to be detected into the trained marine organism detection model to obtain the marine organism detection result output by the marine organism detection model.
[0164] It should be noted that the marine organism detection method provided in the second embodiment of the present application and the marine organism detection model training method in the first embodiment of the present application belong to the same concept. The implementation process is detailed in the method embodiment and will not be repeated here.
[0165] The device embodiments described above are merely illustrative. The components described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present application. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0166] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can be implemented in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can be implemented in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0167] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the selected functions in the process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the selected functions in multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the selected functions in the process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the selected functions in multiple blocks.
[0168] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the selected functions in the process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the selected functions in multiple blocks.
[0169] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0170] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0171] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0172] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0173] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.
Claims
1. A marine organism detection model training method, characterized in that: Applied to a marine life detection model to be trained; the marine life detection model comprises a backbone network, a neck network and a head network connected in sequence; the backbone network comprises a convolutional layer, a plurality of convolutional attention layers and a spatial pooling layer; Methods include: Inputting the marine organism training sample into the convolution layer for feature extraction to obtain a first sample feature; Inputting the first sample feature into the plurality of cascaded convolutional attention layers for feature extraction and feature attention processing to obtain second sample features output by each convolutional attention layer; Inputting the second sample feature output by the last convolutional attention layer into the spatial pooling layer for dimensionality reduction processing to obtain a pooling feature; Inputting the pooled features and a plurality of preset second sample features into the neck network for fusion processing and spatial transformation attention processing to obtain a plurality of fused attention features; The multiple fused attention features are input into the head network for predictive learning processing to obtain a trained marine life detection model.
2. The marine organism detection model training method according to claim 1, characterized in that: The convolutional attention layer includes a convolution module and a dual attention module; The convolution module performs feature extraction processing on the input to obtain a third sample feature; The dual attention module performs feature attention processing on the third sample feature of the same convolutional attention layer to obtain the second sample feature, and the dual attention module outputs the second sample feature to the convolution module of the next-level convolutional attention layer.
3. The marine organism detection model training method according to claim 2, characterized in that: The dual attention module includes an enhanced convolution block, a dual adaptation attention block and a first superposition fusion block; The enhanced convolution block performs convolution enhancement processing on the third sample feature to obtain a fourth sample feature; The dual adaptive attention block performs feature attention fusion processing on the third sample feature to obtain a fifth sample feature; The first superposition and fusion block performs superposition and fusion processing on the third sample feature, the fourth sample feature and the fifth sample feature to obtain the second sample feature.
4. The marine organism detection model training method according to claim 3, characterized in that: The dual-adaptive attention block includes: a first convolution activation submodule, a second convolution activation submodule, a decomposition attention fusion submodule and a feature fusion submodule; The first convolution activation submodule performs convolution activation processing on the third sample feature to obtain a first activation feature; The second convolution activation submodule performs convolution activation processing on the first activation feature to obtain a second activation feature; The decomposition and attention fusion submodule performs decomposition processing and attention fusion convolution processing on the first activation feature and the second activation feature to obtain an attention fusion feature; The feature fusion submodule fuses the first activation feature, the second activation feature and the attention fusion feature to obtain the fifth sample feature.
5. The marine organism detection model training method according to claim 4, characterized in that: The decomposition attention fusion submodule includes a decomposition unit, a fusion unit, an attention normalization processing unit and an attention fusion convolution unit; The decomposition unit is used to decompose the first activation feature and the second activation feature to obtain a first decomposition feature and a second decomposition feature; The fusion unit fuses the first decomposition feature and the second decomposition feature to obtain a decomposition fusion feature; The attention normalization processing unit performs attention processing and normalization processing on the decomposed fusion features to obtain attention normalized features; The attention fusion convolution unit performs fusion convolution processing on the first decomposition feature, the second decomposition feature and the attention normalization feature to obtain the attention fusion feature.
6. The marine organism detection model training method according to claim 1, characterized in that: The spatial pooling layer includes a lightweight convolution activation module, a deformation attention module, a pooling fusion module and a dilated convolution module; The lightweight convolution activation module performs convolution processing on the second sample feature output by the last convolution attention layer to obtain a sample activation feature, and the lightweight convolution activation module outputs the sample activation feature to the deformation attention module and the pooling fusion module respectively; The deformation attention module performs offset deformation processing, feature attention processing and fusion processing on the sample activation feature to obtain a sixth sample feature, and the deformation attention module outputs the sixth sample feature to the pooling fusion module; The pooling fusion model performs superposition fusion and pooling fusion processing on the sample activation feature and the sixth sample feature to obtain a first pooling fusion feature; The dilated convolution module performs dilated convolution processing on the first pooled fusion features to obtain the pooled features.
7. The marine organism detection model training method according to claim 6, characterized in that: The deformation attention module includes a first attention parameter acquisition submodule, a deformation submodule, a second attention parameter acquisition submodule and a deformation attention fusion submodule; The first attention parameter acquisition submodule extracts attention parameters from the sample activation features to obtain a first attention parameter; The deformation submodule performs an offset deformation process on the sample activation feature according to the first attention parameter to obtain an offset deformation feature; The second attention parameter acquisition submodule processes the offset deformation feature and the first attention parameter; Perform attention parameter extraction to obtain a second attention parameter and a third attention parameter; The deformed attention fusion submodule fuses the first attention parameter, the second attention parameter and the third attention parameter to obtain the sixth sample feature.
8. The marine organism detection model training method according to claim 1, characterized in that: The neck network includes a first upsampling fusion module, a first transformation attention module, a second upsampling fusion module, a second transformation attention module, a first convolution fusion module, a third transformation attention module, a second convolution fusion module and a fourth transformation attention module; After the first upsampling fusion module performs upsampling processing on the pooled feature, it fuses the upsampling with the second sample feature output by the penultimate convolutional attention layer to obtain a first upsampling fusion feature; The first transformation attention module performs a transformation attention process on the first up-sampled fusion feature to obtain a first transformation attention feature; The second upsampling fusion module performs upsampling processing on the first transformed attention feature, and then fuses it with the second sample feature output by the third-to-last convolutional attention layer to obtain a second upsampling fusion feature; The second transformed attention module performs a transformed attention process on the first transformed attention feature to obtain a second transformed attention feature; The first convolution fusion module performs convolution processing on the second transformed attention feature and then fuses it with the first transformed attention feature to obtain a first convolution fusion feature; The third change attention module performs change attention processing on the first convolution fusion feature to obtain a third change attention feature; The second convolution fusion module performs a change attention process on the three-transformation attention feature and fuses it with the pooling feature to obtain a second convolution fusion feature; The fourth change attention module performs change attention processing on the second convolution fusion feature to obtain a fourth change attention feature; The second transformed attention feature, the third transformed attention feature and the fourth transformed attention feature are determined as the fused attention feature.
9. The marine organism detection model training method according to claim 8, characterized in that: The enhanced convolution block performs convolution enhancement processing on the third sample feature to obtain a fourth sample feature, including: ecb_output=γ ECB ·ECB(BN(X A )) ecb_output is the fourth sample feature; X A is the third sample feature; BN(·) is the normalization operation; ECB(·) is the convolution enhancement operation; γ ECB Learning parameters for convolutional augmentation processing; The dual adaptive attention block performs feature attention fusion processing on the third sample feature to obtain a fifth sample feature, including: attention_output=γ DAA ·DAA(BN(X A )) Among them, attention_output is the fifth sample feature; DAA(·) is the feature attention fusion operation; γ DAA is the learning parameter for feature attention fusion; The first superposition and fusion block performs superposition and fusion processing on the third sample feature, the fourth sample feature and the fifth sample feature to obtain the second sample feature, including: Y A =attention_output+ecb_output+X A Among them, Y A is the second sample feature.
10. A method for detecting marine organisms, characterized in that: include: Obtaining a trained marine organism detection model by the marine organism detection model training method according to any one of claims 1 to 9; The marine image to be detected is input into the trained marine life detection model to obtain the marine life detection result output by the marine life detection model.
Citation Information
Patent Citations
Scene Tibetan detection method based on improved double-attention YOLOv7
CN116246282A
Fall detection method based on lightweight YOLOv8 network
CN117912111A
Face detection method and device based on YOLOv8 target detection model
CN118470767A
Eel fry detecting and counting method
CN118898859A
Remote sensing vehicle target detection method based on feature fusion and attention mechanism
CN119274037A