Underwater detection method, device, medium and electronic equipment based on underwater acoustic and optical characteristics

By fusing optical and acoustic features in underwater detection, and using contrast learning and channel attention fusion methods, the problem of missing texture information of underwater optical image is solved, significantly improving the accuracy of underwater target detection.

CN117710803BActive Publication Date: 2025-05-16INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410053431.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-12
Publication Date
2025-05-16
Estimated Expiration
2044-01-12

AI Technical Summary

Technical Problem

The lack of texture information of underwater optical image has led to the underwater object detection task posing a major challenge to the general object detection algorithm.

Method used

Pre-box features are generated for optical and acoustic images using two class-independent detectors, and the final detection results are obtained by comparing learning and channel attention fusion methods, combining the RCNN network for pre-box classification and border regression.

Benefits of technology

The texture characteristics of the underwater target are greatly supplemented, significantly improving the target detection accuracy of underwater blurred optical images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117710803B_ABST
    Figure CN117710803B_ABST
Patent Text Reader

Abstract

The present invention relates to an underwater detection method based on underwater acoustic and optical features, comprising: obtaining acoustic detector information; obtaining acoustic pre-selection frame features based on the acoustic detector information; obtaining optical image detector information; obtaining optical pre-selection frame features based on the optical image detector information; dimensional splicing of acoustic pre-selection frame features and dimensional splicing of optical pre-selection frame features; fusing the dimensionally spliced ​​acoustic pre-selection frame features and the dimensionally spliced ​​optical pre-selection frame features to generate fused pre-selection frame features; and determining detection results based on the fused pre-selection frame features. The texture features of underwater targets are greatly supplemented, and the detection results are significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to an underwater detection method, device, medium and electronic equipment based on underwater sound and light characteristics. Background Art

[0002] Target detection in underwater environments is of great value for the development of marine resources, protection of the marine environment, and marine search and rescue. For example, for marine search and rescue, accurate and rapid detection of search and rescue targets is crucial for disaster relief; accurate detection and counting of marine species can provide important data for marine species protection. The underwater imaging environment has its own particularity. The absorption and scattering of light by water will cause the color of underwater images to deflect, reduce the contrast of the image, and make the image blurred. These factors will further lead to the loss of texture information of underwater optical images. In the absence of texture information, the underwater target detection task will pose a major challenge to general target detection algorithms.

[0003] The problem of missing texture in underwater optical images can be compensated by supplementing image features of additional modes. Since the attenuation of sound waves in seawater is much smaller than that of electromagnetic waves, sound waves are an effective medium for detecting targets and transmitting information in the sea. Therefore, sonar based on the characteristics of sound waves has become an important tool for underwater detection. Considering the special properties of sonar underwater, underwater sonar images can have higher imaging quality. Therefore, additional sonar images can be used to fuse underwater optical images to compensate for the problem of missing texture information in underwater optical images, thereby improving the target detection accuracy of underwater blurred optical images. Based on this, this patent proposes a target detection method that integrates underwater acoustic and optical features.

[0004] The present invention proposes an underwater detection method based on underwater acoustic and optical features. First, two category-independent detectors are used to generate pre-selected boxes that may contain targets of interest for optical images and acoustic images, respectively, and the features of each pre-selected box are obtained using a pooling method. Then, based on the idea of ​​contrastive learning, the similarity matrix of the optical pre-selected box features and the acoustic pre-selected box features is calculated, and a one-to-one match between the optical feature pre-selected box and the acoustic feature pre-selected box is established based on this matrix. Secondly, a fusion method is designed using the idea of ​​channel attention to achieve effective fusion of paired acoustic features and optical features. Finally, the fused features are input into a category-related detection network such as an RCNN network to perform pre-selected box classification and border regression to obtain the final detection result. The texture features of underwater targets are greatly supplemented, and the detection results are significantly improved. Summary of the invention

[0005] In view of this, the present invention proposes an underwater detection method based on underwater acoustic and optical features. First, two category-independent detectors are used to generate pre-selected boxes that may contain targets of interest for optical images and acoustic images, respectively, and the features of each pre-selected box are obtained using a pooling method. Then, based on the idea of ​​contrastive learning, the similarity matrix of the optical pre-selected box features and the acoustic pre-selected box features is calculated, and a one-to-one match between the optical feature pre-selected box and the acoustic feature pre-selected box is established based on this matrix. Secondly, a fusion method is designed using the idea of ​​channel attention to achieve effective fusion of paired acoustic features and optical features. Finally, the fused features are input into a category-related detection network such as an RCNN network to perform pre-selected box classification and border regression to obtain the final detection result.

[0006] Specifically, the present invention is achieved through the following technical solutions:

[0007] According to a first aspect of the present invention, there is provided an underwater detection method based on underwater acoustic and optical features, comprising: obtaining acoustic detector information; obtaining acoustic pre-selection box features based on the acoustic detector information; obtaining optical image detector information; obtaining optical pre-selection box features based on the optical image detector information; dimensional splicing of the acoustic pre-selection box features and dimensional splicing of the optical pre-selection box features; fusing the dimensionally spliced ​​acoustic pre-selection box features and the dimensionally spliced ​​optical pre-selection box features to generate fused pre-selection box features; and obtaining detection results based on the fused pre-selection box features.

[0008] In the above technical solution, the step of obtaining acoustic pre-selection box features based on acoustic detector information includes:

[0009] The acoustic pre-selection box is obtained in the acoustic detector information based on the YOLO series algorithm, wherein the YOLO series algorithm is as follows:

[0010]

[0011] Among them, YOLO represents the category-independent YOLO series target detection algorithm, A represents the input acoustic image, represents the i-th pre-selected box detected in the acoustic image, P a represents the features extracted from the acoustic image;

[0012] The step of acquiring an optical pre-selection frame based on the optical image detector information comprises:

[0013] Acquiring the acoustic pre-selection box in the optical detector information based on the YOLO series algorithm;

[0014]

[0015] Among them, YOLO represents the category-independent YOLO series target detection algorithm, O represents the input optical image, represents the i-th pre-selected box detected in the optical image, P o Represents the features extracted from the optical image.

[0016] In the above technical solution, before the steps of dimensional splicing of acoustic pre-selection frame features and dimensional splicing of optical pre-selection frame features, the following steps are also included:

[0017] Pooling is performed on the acoustic pre-selected box features and the optical pre-selected box features.

[0018] In the above technical solution, the step of performing pooling processing on the acoustic pre-selection box features and the optical pre-selection box features includes:

[0019] The acoustic pre-selection box features and the optical pre-selection box features are pooled based on a Pooling algorithm.

[0020]

[0021]

[0022] Among them, Pooling represents the pooling operation. represents the feature obtained by pooling the i-th pre-selected box in the optical image, Represents the features obtained by pooling the i-th pre-selected box in the acoustic image.

[0023] In the above technical solution, the formula for dimensional splicing of the optical pre-selection frame feature is as follows:

[0024]

[0025] Among them, Π represents the splicing operation, F o represents the optical pre-selected box feature after splicing, N represents the number of pre-selected boxes generated in the optical image, Represents the features obtained by pooling the i-th pre-selected box in the optical image;

[0026] The formula for dimensional concatenation of the acoustic pre-selected box features is as follows:

[0027]

[0028] Among them, Π represents the splicing operation, F a represents the acoustic pre-selected box features after splicing, M represents the number of pre-selected boxes generated in the acoustic image, Represents the features obtained by pooling the i-th pre-selected box in the acoustic image.

[0029] In the above technical solution, after the steps of dimensional splicing of acoustic pre-selected box features and dimensional splicing of optical pre-selected box features, the method further includes: calculating a similarity matrix based on matrix multiplication and a softmax function, and the specific formula is as follows;

[0030] S oa =Softmax(F o ·F aT ),

[0031] Where Softmax represents the activation function, S oa represents the calculated similarity matrix, The transposed matrix representing the concatenated acoustic pre-selected box features;

[0032] The traction of the acoustic pre-selection box features corresponding to the optical pre-selection box features is obtained based on the similarity matrix. The specific formula is as follows;

[0033] I oa =ArgMax(S oa ),

[0034] Among them, ArgMax is the function for finding the maximum value index, I oa It is the index of the matching of the optical pre-selected box feature to the acoustic pre-selected box feature.

[0035] In the above technical solution, the step of fusing the acoustic pre-selection frame features after dimension splicing and the optical pre-selection frame features after dimension splicing includes:

[0036] Use a multi-layer perceptron to construct the input of channel attention. The specific formula is as follows:

[0037]

[0038] Among them, MLP is a multi-layer perceptron. is the input of channel attention, Represents the features obtained by pooling the jth pre-selected box in the acoustic image;

[0039] The input of the channel attention is fused to obtain the fusion weight of the optical pre-selection box feature and the fusion weight of the acoustic pre-selection box feature. The specific formula is as follows:

[0040]

[0041] Among them, AvgPool and MaxPool are maximum pooling functions, and Split represents a splitting function, which is used to split the features into two equal parts in the channel dimension. is the assigned fusion weight of the optical pre-selected box feature, is the assigned fusion weight of the acoustic pre-selected box feature;

[0042] The acoustic pre-selection box features after dimension splicing and the optical pre-selection box features after dimension splicing are fused through the fusion weights assigned to the optical pre-selection box features and the fusion weights assigned to the acoustic pre-selection box features to generate a fused pre-selection box. The specific formula is as follows:

[0043]

[0044] in, Indicates the fusion of pre-selected boxes, and * indicates the weighted summation.

[0045] In the above technical solution, the step of obtaining the detection result based on the fusion of the pre-selected box features includes:

[0046] Use the RCNN network to process the fused pre-selected box features and generate the detection results. The specific formula is as follows;

[0047]

[0048] Among them B ij Represents the coordinates of the predicted bounding box, C ij Indicates the predicted category.

[0049] According to a second aspect of the present invention, an underwater detection device based on underwater acoustic and optical features is provided, the device comprising a module for executing the underwater detection method based on underwater acoustic and optical features in the first aspect or any possible implementation of the first aspect.

[0050] According to a third aspect of the present invention, there is provided a storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the program implements the steps of the graph matching method based on reinforcement learning in the first aspect or any possible implementation of the first aspect.

[0051] According to a fourth aspect of the present invention, there is provided an electronic device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the graph matching method based on reinforcement learning in the first aspect or any possible implementation of the first aspect are implemented.

[0052] The technical solution provided by the present invention brings at least the following beneficial effects: First, an optical detector and an acoustic detector are used to generate pre-selected box features for optical images and acoustic images respectively, and a pooling method is used to obtain the features of each pre-selected box feature. Then, based on the idea of ​​contrastive learning, the similarity matrix of the optical pre-selected box features and the acoustic pre-selected box features is calculated, and a one-to-one match between the optical feature pre-selected box features and the acoustic feature pre-selected box features is established based on this matrix. Secondly, a fusion method is designed using the idea of ​​channel attention to achieve effective fusion of paired acoustic features and optical features. Finally, the fused features are input into a category-related detection network such as an RCNN network to perform classification and border regression of the pre-selected box features to obtain the final detection result. The texture features of underwater targets are greatly supplemented, and the detection results are significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0054] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or related technical descriptions are briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0055] Figure 1 A schematic diagram of the process of underwater detection based on underwater acoustic and optical characteristics provided by an embodiment of the present invention;

[0056] Figure 2 A block diagram of an underwater detection device based on underwater acoustic and optical characteristics provided by an embodiment of the present invention;

[0057] Figure 3 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention.

[0058] 200 underwater detection device based on underwater acoustic and optical characteristics, 201 acoustic detector, 202 optical detector, 203 pre-selection box feature acquisition device, 204 dimension stitching device, 205 pre-selection box feature fusion device, 206 result acquisition device, 3 electronic device, 31 memory, 32 processor. DETAILED DESCRIPTION

[0059] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0060] like Figure 1 As shown, the embodiment of the present invention provides an underwater detection method based on underwater acoustic and optical characteristics, including:

[0061] S101: Acquiring acoustic detector information;

[0062] S102: Acquiring acoustic pre-selection box features based on acoustic detector information;

[0063] S103: Acquire optical image detector information;

[0064] S104: Acquire optical pre-selection box features based on optical image detector information;

[0065] S105: performing dimensional splicing on acoustic pre-selected box features and performing dimensional splicing on optical pre-selected box features;

[0066] S106: Fusing the acoustic pre-selected box features after dimension splicing and the optical pre-selected box features after dimension splicing to generate a fused pre-selected box feature;

[0067] S107: Determine the detection result based on the fused pre-selected box features.

[0068] In the above technical solution, the step of obtaining acoustic pre-selection box features based on acoustic detector information includes:

[0069] The acoustic pre-selection box is obtained in the acoustic detector information based on the YOLO series algorithm, wherein the YOLO series algorithm is as follows:

[0070]

[0071] Among them, YOLO represents the category-independent YOLO series target detection algorithm, A represents the input acoustic image, represents the i-th pre-selected box detected in the acoustic image, P a represents the features extracted from the acoustic image;

[0072] The step of acquiring an optical pre-selection frame based on the optical image detector information comprises:

[0073] Acquiring the acoustic pre-selection box in the optical detector information based on the YOLO series algorithm;

[0074]

[0075] Among them, YOLO represents the category-independent YOLO series target detection algorithm, O represents the input optical image, represents the i-th pre-selected box detected in the optical image, P o Represents the features extracted from the optical image.

[0076] The conventional YOLO algorithm will simultaneously perform target detection and determine the category of an object, while the YOLO algorithm used in this application will not determine the category of the object and will only output the target detection result.

[0077] In the above technical solution, before the steps of dimensional splicing of acoustic pre-selection frame features and dimensional splicing of optical pre-selection frame features, the following steps are also included:

[0078] Pooling is performed on the acoustic pre-selected box features and the optical pre-selected box features.

[0079] In the above technical solution, the step of performing pooling processing on the acoustic pre-selection box features and the optical pre-selection box features includes:

[0080] The acoustic pre-selection box features and the optical pre-selection box features are pooled based on a Pooling algorithm.

[0081]

[0082]

[0083] Among them, Pooling represents the pooling operation. represents the feature obtained by pooling the i-th pre-selected box in the optical image, Represents the features obtained by pooling the i-th pre-selected box in the acoustic image.

[0084] In the above technical solution, the formula for dimensional splicing of the optical pre-selection frame feature is as follows:

[0085]

[0086] Among them, Π represents the splicing operation, F o represents the optical pre-selected box feature after splicing, N represents the number of pre-selected boxes generated in the optical image, Represents the features obtained by pooling the i-th pre-selected box in the optical image;

[0087] The formula for dimensional concatenation of the acoustic pre-selected box features is as follows:

[0088]

[0089] Among them, Π represents the splicing operation, F a represents the acoustic pre-selected box features after splicing, M represents the number of pre-selected boxes generated in the acoustic image, Represents the features obtained by pooling the i-th pre-selected box in the acoustic image.

[0090] In the above technical solution, after the steps of dimensional splicing of acoustic pre-selected box features and dimensional splicing of optical pre-selected box features, the method further includes: calculating a similarity matrix based on matrix multiplication and a softmax function, and the specific formula is as follows;

[0091]

[0092] Where Softmax represents the activation function, S oa represents the calculated similarity matrix, The transposed matrix representing the concatenated acoustic pre-selected box features;

[0093] The traction of the acoustic pre-selection box features corresponding to the optical pre-selection box features is obtained based on the similarity matrix. The specific formula is as follows;

[0094] I oa =ArgMax(S oa ),

[0095] Among them, ArgMax is the function for finding the maximum value index, I oa It is the index of the matching of the optical pre-selected box feature to the acoustic pre-selected box feature.

[0096] In the above technical solution, the step of fusing the acoustic pre-selection frame features after dimension splicing and the optical pre-selection frame features after dimension splicing includes:

[0097] Use a multi-layer perceptron to construct the input of channel attention. The specific formula is as follows:

[0098]

[0099] Among them, MtP is a multi-layer perceptron, is the input of channel attention, Represents the features obtained by pooling the jth pre-selected box in the acoustic image;

[0100] The input of the channel attention is fused to obtain the fusion weight of the optical pre-selection box feature and the fusion weight of the acoustic pre-selection box feature. The specific formula is as follows:

[0101]

[0102] Among them, AvgPool and MaxPool are maximum pooling functions, and Split represents a splitting function, which is used to split the features into two equal parts in the channel dimension. is the assigned fusion weight of the optical pre-selected box feature, is the assigned fusion weight of the acoustic pre-selected box feature;

[0103] The acoustic pre-selection box features after dimension splicing and the optical pre-selection box features after dimension splicing are fused through the fusion weights assigned to the optical pre-selection box features and the fusion weights assigned to the acoustic pre-selection box features to generate a fused pre-selection box. The specific formula is as follows:

[0104]

[0105] in, Indicates the fusion of pre-selected boxes, and * indicates the weighted summation.

[0106] In the above technical solution, the step of obtaining the detection result based on the fusion of the pre-selected box features includes:

[0107] Use the RCNN network to process the fused pre-selected box features and generate the detection results. The specific formula is as follows;

[0108]

[0109] Among them B ij Represents the coordinates of the predicted bounding box, C ij Indicates the predicted category.

[0110] The technical solution provided by the present invention brings at least the following beneficial effects: First, an optical detector and an acoustic detector are used to generate pre-selected box features for optical images and acoustic images respectively, and a pooling method is used to obtain the features of each pre-selected box feature. Then, based on the idea of ​​contrastive learning, the similarity matrix of the optical pre-selected box features and the acoustic pre-selected box features is calculated, and a one-to-one match between the optical feature pre-selected box features and the acoustic feature pre-selected box features is established based on this matrix. Secondly, a fusion method is designed using the idea of ​​channel attention to achieve effective fusion of paired acoustic features and optical features. Finally, the fused features are input into a category-related detection network such as an RCNN network to perform classification and border regression of the pre-selected box features to obtain the final detection result. The texture features of underwater targets are greatly supplemented, and the detection results are significantly improved.

[0111] We collected and annotated 5,000 underwater sound and light image pairs (including 10 common marine species categories) to verify the method of this application. We divided these 5,000 samples into 3,000 training samples, 1,000 validation samples, and 1,000 test samples. We used YOLOv6, YOLOv7, and YOLOv8 as baseline methods, and the experimental results are shown in Table 1. The algorithm proposed in this invention significantly surpasses the baseline method in all indicators.

[0112] Table 1

[0113]

[0114]

[0115] Based on the same inventive concept, Figure 2 As shown, an embodiment of the present invention further provides an underwater detection device 200 based on underwater acoustic and optical features, the device comprising: an acoustic detector 201, used to obtain acoustic detector information; an optical detector 202, used to obtain optical image detector information; a pre-selection box feature acquisition device 203, used to obtain acoustic pre-selection box features based on the acoustic detector information and / or to obtain optical pre-selection box features based on the optical image detector information; a dimension splicing device 204, used to dimensionally splice the acoustic pre-selection box features and the optical pre-selection box features; a pre-selection box feature fusion device 205, used to fuse the acoustic pre-selection box features after dimension splicing and the optical pre-selection box features after dimension splicing to generate a fused pre-selection box feature; a result acquisition device 206, used to obtain a detection result based on the fused pre-selection box feature.

[0116] In the above embodiment, first, the acoustic detector information and the optical image detector information are obtained through the acoustic detector 201 and the optical detector 202. The pre-selection box feature acquisition device 203 acquires the acoustic pre-selection box feature based on the acoustic detector information and acquires the optical pre-selection box feature based on the optical image detector information. The dimension splicing device 204 performs dimension splicing on the acoustic pre-selection box feature and dimension splicing on the optical pre-selection box feature. The pre-selection box feature fusion device 205 fuses the dimensionally spliced ​​acoustic pre-selection box feature and the dimensionally spliced ​​optical pre-selection box feature to generate a fused pre-selection box feature; the result acquisition device 206 acquires the detection result based on the fused pre-selection box feature.

[0117] The implementation process of the functions and effects of each module in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, which will not be repeated here.

[0118] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of the present invention. Ordinary technicians in this field can understand and implement it without paying creative work.

[0119] Based on the same inventive concept, an embodiment of the present invention further provides a storage medium on which a computer program is stored, and when the program is executed by a processor, the steps of the graph matching method based on reinforcement learning in any possible implementation manner described above are implemented.

[0120] Alternatively, the storage medium may be a non-transitory computer-readable storage medium, for example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, and the like.

[0121] Based on the same inventive concept, see Figure 3 The embodiment of the present invention further provides an electronic device 3, including a memory 31 (such as a non-volatile memory), a processor 32, and a computer program stored in the memory 31 and executable on the processor 32. When the processor 32 executes the program, the steps of the graph matching method based on reinforcement learning in any possible implementation manner described above are implemented, which may be equivalent to the graph matching device based on reinforcement learning as described above. Of course, the processor may also be used to process other data or operations. The electronic device may be a PC, a server, a terminal, or other equipment.

[0122] The embodiments of the subject matter and functional operations described in this specification may be implemented in the following: digital electronic circuits, tangibly embodied computer software or firmware, computer hardware including the structures disclosed in this specification and their structural equivalents, or a combination of one or more of them. The embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules in computer program instructions encoded on a tangible non-temporary program carrier to be executed by a data processing device or to control the operation of the data processing device. Alternatively or additionally, the program instructions may be encoded on an artificially generated propagation signal, such as a machine-generated electrical, optical or electromagnetic signal, which is generated to encode information and transmit it to a suitable receiver device for execution by a data processing device. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.

[0123] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform corresponding functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuits, such as FPGAs (field programmable gate arrays) or ASICs (application-specific integrated circuits), and the apparatus can also be implemented as special purpose logic circuits.

[0124] Computers suitable for executing computer programs include, for example, general and / or special microprocessors, or any other type of central processing unit. Typically, the central processing unit will receive instructions and data from a read-only memory and / or a random access memory. The basic components of a computer include a central processing unit for implementing or executing instructions and one or more memory devices for storing instructions and data. Typically, the computer will also include one or more large-capacity storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or the computer will be operably coupled to this large-capacity storage device to receive data from it or to transmit data to it, or both. However, the computer does not necessarily have such a device. In addition, the computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name a few.

[0125] Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special purpose logic circuitry.

[0126] Although this specification includes many specific implementation details, these should not be interpreted as limiting the scope of any invention or the scope of protection claimed, but are mainly used to describe the features of the specific embodiments of specific inventions. Certain features described in multiple embodiments in this specification may also be implemented in combination in a single embodiment. On the other hand, the various features described in a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. In addition, although features may work in certain combinations as described above and even initially claim protection, one or more features from the claimed combination may be removed from the combination in some cases, and the claimed combination may point to a sub-combination or a variation of a sub-combination.

[0127] Similarly, although operations are depicted in a particular order in the accompanying drawings, this should not be understood as requiring that these operations be performed in the particular order shown or performed sequentially, or requiring that all illustrated operations be performed to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of various system modules and components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product, or packaged into multiple software products.

[0128] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired results. In some implementations, multitasking and parallel processing may be advantageous.

[0129] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0130] The foregoing is merely a specific embodiment of the present invention, which enables those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features claimed herein.

Claims

1. An underwater detection method based on underwater acoustic and optical characteristics, characterized in that: include: Get acoustic detector information; Acquiring acoustic pre-selection box features based on the acoustic detector information; Acquiring optical image detector information; Acquire optical pre-selection box features based on the optical image detector information; Performing dimension stitching on the acoustic pre-selection frame features and performing dimension stitching on the optical pre-selection frame features; Fusing the acoustic pre-selection frame features after dimension splicing and the optical pre-selection frame features after dimension splicing to generate a fused pre-selection frame feature; Determine a detection result based on the fused pre-selected box feature; The formula for dimensional splicing of the optical pre-selection frame feature is as follows: Among them, Π represents the splicing operation, F o represents the optical pre-selected box feature after splicing, N represents the number of pre-selected boxes generated in the optical image, Represents the features obtained by pooling the i-th pre-selected box in the optical image; The formula for dimensional concatenation of the acoustic pre-selected box features is as follows: Among them, Π represents the splicing operation, F a represents the acoustic pre-selected box features after splicing, M represents the number of pre-selected boxes generated in the acoustic image, Represents the features obtained by pooling the i-th pre-selected box in the acoustic image; After the steps of dimensional splicing of the acoustic pre-selection box features and dimensional splicing of the optical pre-selection box features, the method further includes: calculating a similarity matrix based on matrix multiplication and a softmax function, and the specific formula is as follows; Among them, softmax represents the activation function, S oa represents the calculated similarity matrix, The transposed matrix representing the concatenated acoustic pre-selected box features; Based on the similarity matrix, the index of the acoustic pre-selection box feature corresponding to the optical pre-selection box feature is obtained, and the specific formula is as follows: I oa =ArgMax(S oa ); Among them, ArgMax is the function for finding the maximum value index, I oa It is the index of the matching of the optical pre-selected box feature to the acoustic pre-selected box feature.

2. The underwater detection method based on underwater acoustic and optical characteristics according to claim 1 is characterized in that: The step of acquiring acoustic pre-selection box features based on the acoustic detector information comprises: The acoustic pre-selection box is obtained in the acoustic detector information based on the YOLO series algorithm, wherein the YOLO series algorithm is as follows: Among them, YOLO represents the category-independent YOLO series target detection algorithm, A represents the input acoustic image, represents the i-th pre-selected box detected in the acoustic image, P a represents the features extracted from the acoustic image; The step of acquiring an optical pre-selection frame based on the optical image detector information comprises: Acquire the optical pre-selection box in the optical image detector information based on the YOLO series algorithm; Among them, YOLO represents the category-independent YOLO series target detection algorithm, O represents the input optical image, represents the i-th pre-selected box detected in the optical image, P o Represents the features extracted from the optical image.

3. The underwater detection method based on underwater acoustic and optical characteristics according to claim 2 is characterized in that: Before the step of dimensionally splicing the acoustic pre-selection frame feature and the optical pre-selection frame feature, the step further includes: Pooling is performed on the acoustic pre-selection box features and the optical pre-selection box features.

4. The underwater detection method based on underwater acoustic and optical characteristics according to claim 3 is characterized in that: The step of performing pooling processing on the acoustic pre-selection box features and the optical pre-selection box features comprises: The acoustic pre-selection box features and the optical pre-selection box features are pooled based on a Pooling algorithm. Among them, Pooling represents the pooling operation. represents the feature obtained by pooling the i-th pre-selected box in the optical image, Represents the features obtained by pooling the i-th pre-selected box in the acoustic image.

5. The underwater detection method based on underwater acoustic and optical characteristics according to claim 4 is characterized in that: The step of fusing the dimensionally spliced ​​acoustic pre-selection frame features and the dimensionally spliced ​​optical pre-selection frame features comprises: Use a multi-layer perceptron to construct the input of channel attention. The specific formula is as follows: Among them, MLP is a multi-layer perceptron. is the input of channel attention, Represents the features obtained by pooling the jth pre-selected box in the acoustic image; The input of the channel attention is fused to obtain the fusion weight of the optical pre-selection box feature and the fusion weight of the acoustic pre-selection box feature. The specific formula is as follows: Among them, AvgPool represents the average pooling function, MaxPool represents the maximum pooling function, and Split represents the splitting function, which is used to split the feature into two equal parts in the channel dimension. is the assigned fusion weight of the optical pre-selected box feature, is the assigned fusion weight of the acoustic pre-selected box feature; The acoustic pre-selection box features after dimension splicing and the optical pre-selection box features after dimension splicing are fused through the fusion weights assigned to the optical pre-selection box features and the fusion weights assigned to the acoustic pre-selection box features to generate a fused pre-selection box. The specific formula is as follows: in, Indicates the fusion of pre-selected boxes, and * indicates the weighted summation.

6. The underwater detection method based on underwater acoustic and optical characteristics according to claim 5 is characterized in that: The step of determining the detection result based on the fused pre-selected box feature comprises: Using the RCNN network to process the fused pre-selected box features to generate a detection result; Among them, B ij Represents the coordinates of the predicted bounding box, C ij Indicates the predicted category.

7. An underwater detection device based on underwater acoustic and optical characteristics, characterized in that: include: An acoustic detector, used for obtaining acoustic detector information; An optical detector, used for acquiring optical image detector information; A pre-selection box acquisition device, used for acquiring acoustic pre-selection box features based on the acoustic detector information and acquiring optical pre-selection box features based on the optical image detector information; A dimensional splicing device, used for dimensional splicing the acoustic pre-selection frame features and dimensional splicing the optical pre-selection frame features; A pre-selection frame fusion device, used for fusing the acoustic pre-selection frame features after dimension splicing and the optical pre-selection frame features after dimension splicing to generate a fused pre-selection frame feature; A result acquisition device, used to determine the detection result based on the fused pre-selected box feature; The formula for dimensional concatenation of the acoustic pre-selected box features is as follows: Among them, Π represents the splicing operation, F o represents the optical pre-selected box features after splicing, and N represents the number of pre-selected boxes generated in the acoustic image; The formula for dimensional splicing of the optical pre-selection frame feature is as follows: Among them, Π represents the splicing operation, F a represents the acoustic pre-selected box features after splicing, and M represents the number of pre-selected boxes generated in the optical image; The pre-selected box fusion device is used to calculate a similarity matrix based on matrix multiplication and softmax function, and the specific formula is as follows; Among them, softmax represents the activation function, S oa represents the calculated similarity matrix, The transposed matrix representing the concatenated acoustic pre-selected box features; Based on the similarity matrix, the index of the acoustic pre-selection box feature corresponding to the optical pre-selection box feature is obtained, and the specific formula is as follows: I oa =ArgMax(S oa ); Among them, ArgMax is the function for finding the maximum value index, I oa It is the index of the matching of the optical pre-selected box feature to the acoustic pre-selected box feature.

8. A storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Fish behavior identification method based on multi-level fusion of sound and vision

    CN115170942A