Submersible data acquisition method and device based on deep learning
By employing a deep learning-based data acquisition method for submersibles, and utilizing historical image data and manually labeled models to train the model, autonomous target recognition and data acquisition by the submersible in the underwater environment were achieved. This solved the problem of low efficiency in traditional methods and improved the real-time performance and accuracy of data acquisition.
Patent Information
- Application Number
- CN202511461773.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-10-14
AI Technical Summary
Traditional methods rely on manual screening and labeling of submersible images, which is inefficient and cannot meet real-time requirements, resulting in low underwater image quality and difficulty in target identification.
A deep learning-based approach is adopted to train a deep learning model by acquiring historical image data and manually labeled data, thereby constructing an underwater target recognition model. Image data is collected and enhanced in real time during submersible operation, and the model is used for target recognition, control of data acquisition and storage.
It enables submersibles to autonomously complete target discovery and data acquisition in complex underwater environments, reducing reliance on manual operation and remote communication, expanding the scope and duration of operations, and improving the efficiency and accuracy of data acquisition.
Smart Images

Figure CN120932083A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of underwater exploration and artificial intelligence technology, and more specifically, to a deep learning-based method and apparatus for data acquisition from a submersible. Background Technology
[0002] Submersibles typically need to collect large amounts of image or video data during underwater operations for tasks such as environmental monitoring, target identification, and resource exploration. However, due to the complex and variable underwater environment, poor lighting conditions, and severe water scattering, the quality of the acquired images is low, making target identification difficult. Traditional methods often rely on manual image screening and labeling, which is not only inefficient but also fails to meet real-time requirements. Summary of the Invention
[0003] This application aims to provide a deep learning-based method and apparatus for underwater vehicle data acquisition, which addresses the problems of traditional methods relying on manual image screening and labeling, which are not only inefficient but also unable to meet real-time requirements.
[0004] The first aspect of this application provides a deep learning-based submersible data acquisition method, comprising: Acquire historical image data collected by the submersible at historical moments and manually labeled tags corresponding to the historical image data; wherein, the historical image data is data after image enhancement processing; A deep learning model is constructed and trained based on the historical image data and manually labeled information corresponding to the historical image data to obtain an underwater target recognition model. During the underwater operation of the submersible, real-time image data is acquired using a preset data sampling frequency, and the real-time image data is subjected to image enhancement processing to obtain enhanced real-time image data. The underwater target recognition model is scheduled to recognize the enhanced real-time image data to determine the target recognition result in the real-time image data; Based on the target recognition results in the real-time image data, the submersible is controlled to collect and store data, thus completing the deep learning-based submersible data collection.
[0005] In one possible design approach, building a deep learning model includes: building a YOLO v3 model to obtain a deep learning model.
[0006] In one possible design approach, the deep learning model is trained based on the historical image data and manually labeled tags corresponding to the historical image data to obtain an underwater target recognition model, including: The hyperparameters of the deep learning model are initialized and encoded to obtain the training population; For any individual in the training population, the fitness of the individual is obtained by training based on the historical image data and manually labeled tags corresponding to the historical image data. Based on the fitness of the individual, determine the first optimal individual in the current training process; Based on the first optimal individual, a normal distribution local development strategy is used to locally develop the individual to obtain the locally developed individual; An information transmission strategy is used to perform information transmission development on the individuals after the partial development, and to obtain the individuals after the information transmission development; A normal distribution global development strategy is adopted to perform global development on the individuals after the information transmission development, and the individuals after global development are obtained. Determine if the current number of training iterations is greater than or equal to the preset maximum number of training iterations. If so, obtain the second optimal individual based on the individuals after global development; otherwise, return to the step of obtaining the first optimal individual. The hyperparameters of the second optimal individual are used as the final hyperparameters of the deep learning model to obtain the underwater target recognition model.
[0007] In one possible design approach, the fitness of an individual is obtained by training based on the historical image data and manually labeled tags corresponding to the historical image data, including: Using the historical image data as input data, and training with manually labeled tags corresponding to the historical image data as the expected output, the loss function value is obtained; The loss function value is added to a preset constant to obtain a non-zero data item, and the reciprocal of the non-zero data item is taken to obtain the fitness of the individual.
[0008] In one possible design approach, based on the first optimal individual, a normal distribution local development strategy is used to locally develop the individual, obtaining the locally developed individual, including: The average position of the training population in the solution space is obtained as follows:
[0009] in, Indicates the first k The first training session i The first individual d The hyperparameter is N, which represents the total number of individuals in the training population. d =1,2,…,D, where D represents the total dimension of hyperparameters in an individual. Indicates the first k The average position during the training cycle.d dimensional hyperparameters; Based on the average position and the first optimal individual, obtain the first... i The simulated mean information for each individual is:
[0010] in, Indicates the first k During the training session i The simulated average information corresponding to each individual d dimensional hyperparameters, The first optimal individual represents the first optimal individual. d dimensional hyperparameters; Based on the average position, the first optimal individual, and the simulated mean information, obtain the first... i The simulated standard deviation information for each individual is as follows:
[0011] in, Indicates the first i The simulated standard deviation information for each individual. d dimensional hyperparameters; The penalty factor is:
[0012] in, Indicates the penalty factor. Represents the first random number between (0,1). This represents the second random number between (0,1). This represents a third random number between (0,1). This represents the fourth random number between (0,1). Represents the cosine function. Represents pi (π). Represents a logarithmic function; Based on the simulated mean information, simulated standard deviation information, and penalty factor, individuals are locally developed to obtain the following individuals after local development:
[0013] in, Indicates the first i The first individual after partial development d Dimensional hyperparameters.
[0014] In one possible design approach, an information delivery strategy is employed to perform information delivery development on the individuals after the partial development, and the individuals after information delivery development are obtained, including: The information transmission vector is:
[0015]
[0016] in, Indicates the first j The information transmission vector corresponding to the individual after partial development is the first d dimensional hyperparameters, Indicates adaptive inertia weights, This represents the fifth random number that is uniformly distributed between (0,1). Represents the sixth random number uniformly distributed between (0,1). Indicates the first k The first training session i After the individual's partial development, the first individual's d dimensional hyperparameters, Indicates the first j The Euclidean distance between an individual after partial development and the first optimal individual; Indicates the maximum number of training iterations. This represents a random number that follows a standard normal distribution and is distributed according to the first normal distribution. This represents the seventh random number between (0,1). This represents the eighth random number between (0, 1); Represents the first individual after random local development. d dimensional hyperparameters; Based on the information transmission vector, the information transmission individual is obtained as follows:
[0017] in, Indicates the first j The information transmission of the individual corresponding to the individual after partial development. d dimensional hyperparameters; Based on the information transmission individual, information transmission development is performed on the individual after the partial development, and the individual after information transmission development is obtained as follows:
[0018] in, Indicates the first j An individual after information transmission development, Indicates the first j The fitness of an individual after partial development. Indicates the first j The fitness of an individual after partial development, corresponding to information transmission. Indicates the first jThe Euclidean distance between an individual after partial development and the individual transmitting information.
[0019] In one possible design approach, a normal distribution global development strategy is employed to globally develop the individuals after the information transmission development, obtaining the globally developed individuals, including: For each individual that has undergone information transmission development, three other individuals that have undergone information transmission development are randomly matched to obtain the first tracking individual, the second tracking individual, and the third tracking individual corresponding to each individual that has undergone information transmission development. Based on the first tracking individual, the second tracking individual, and the third tracking individual, the first tracking vector and the second tracking vector are obtained as follows:
[0020]
[0021] in, Indicates the first k During the training session m The first individual after information transmission development d dimensional hyperparameters, This indicates the first individual being tracked. d dimensional hyperparameters, Represents the first tracking vector's... d dimensional hyperparameters, Indicates the first k During the training session m The fitness of an individual after the development of information transmission. Indicates the fitness of the first tracked individual; The second tracking vector represents the first... d dimensional hyperparameters, The second tracking vector represents the first... d dimensional hyperparameters, The third tracking vector represents the... d dimensional hyperparameters, This indicates the fitness of the second tracked individual. This indicates the fitness of the third tracked individual; Based on the first tracking vector and the second tracking vector, global development is performed on the individual after the information transmission development, and the individual after global development is obtained as follows:
[0022] in, Indicates the first m The first individual after global development d dimensional hyperparameters, Represents the tenth random number between (0,1). This represents a random number that follows a standard normal distribution or a second normal distribution. This represents a random number that follows a standard normal distribution but is distributed in the third normal distribution.
[0023] In one possible design approach, the real-time image data is subjected to image enhancement processing to obtain enhanced real-time image data, including: The real-time image data is subjected to three-channel color compensation, global histogram stretching, and / or HSV space enhancement to obtain enhanced real-time image data.
[0024] In one possible design approach, based on the target recognition results in the real-time image data, the submersible is controlled to perform data acquisition and storage, including: If the target recognition result in the real-time image data contains a preset target, the data sampling frequency is increased, and the increased data sampling frequency is used to collect real-time image data, or the recording function is enabled to collect real-time video data, and the collected real-time image data and real-time video data are stored.
[0025] A second aspect of this application provides a deep learning-based submersible data acquisition device, comprising: The training data acquisition module is used to acquire historical image data collected by the submersible at historical moments and manually labeled tags corresponding to the historical image data; wherein, the historical image data is data after image enhancement processing; A deep learning model is used to construct a deep learning model and train the deep learning model based on the historical image data and manually labeled tags corresponding to the historical image data to obtain an underwater target recognition model. The image enhancement module is used to acquire real-time image data at a preset data sampling frequency during the underwater operation of the submersible, and to perform image enhancement processing on the real-time image data to obtain enhanced real-time image data. The target recognition module is used to schedule the underwater target recognition model to recognize the enhanced real-time image data and determine the target recognition result in the real-time image data. The submersible data acquisition module is used to control the submersible to acquire and store data based on the target recognition results in the real-time image data, thereby completing the deep learning-based submersible data acquisition.
[0026] Beneficial effects: This application provides a deep learning-based submersible data acquisition method and apparatus, which integrates intelligent perception and decision-making control capabilities into the submersible system, giving it a high degree of autonomy. Even in deep-sea environments where communication is limited or impossible, the submersible can still independently complete complex target discovery, tracking, and refined data acquisition tasks. This not only reduces reliance on manual operation and remote communication, lowering operating costs and risks, but also significantly expands the submersible's operational range and duration, enabling it to undertake longer-duration, longer-distance, and more complex ocean exploration missions. Attached Figure Description
[0027] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 This is a flowchart of a deep learning-based submersible data acquisition method proposed in one embodiment of this application; Figure 2 This is a schematic diagram of the structure of a deep learning-based submersible data acquisition device according to an embodiment of this application. Detailed Implementation
[0029] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0030] like Figure 1 As shown in the figure, this application provides a deep learning-based submersible data acquisition method, including: S101. Obtain historical image data collected by the submersible at historical moments and manually labeled tags corresponding to the historical image data; wherein, the historical image data is data after image enhancement processing.
[0031] A large amount of historical image data collected by the submersible in previous missions is acquired, and these images are manually labeled by marine biologists, geologists, or professional annotators. The labels are for targets of interest in the images, such as "coral," "fish A," "shipwreck," and "hydrothermal vent," and their locations (bounding boxes) are marked. Specifically, after acquiring this raw historical image data, it needs to be image augmented before being used as the final training data. This is done to eliminate environmental interference in the original images, unify the data distribution, and improve the effectiveness of subsequent model training. It is worth noting that how to label the training data for the YOLO model is a relatively conventional technique, and this embodiment will not elaborate further.
[0032] S102. Construct a deep learning model and train the deep learning model based on the historical image data and the labels corresponding to the historical image data that are manually labeled, to obtain an underwater target recognition model.
[0033] In this embodiment, the YOLO v3 model, which combines speed and accuracy in the field of target detection, can be selected as the basic deep learning model. After constructing the model framework, the enhanced historical image data and labels prepared in step S101 are used to train the model. The core of the training process is to optimize the hyperparameters of the model to obtain the underwater target recognition model with the best performance.
[0034] S103. During the underwater operation of the submersible, real-time image data is acquired using a preset data sampling frequency, and the real-time image data is subjected to image enhancement processing to obtain enhanced real-time image data.
[0035] When a submersible performs underwater exploration missions, its onboard cameras continuously acquire real-time image data at a preset, low base data sampling frequency (e.g., one image every 2 seconds). Due to the optical characteristics of the underwater environment, the quality of these raw real-time images is usually poor. Therefore, the submersible's image processing unit immediately performs image enhancement processing on each frame of the real-time image.
[0036] S104. The underwater target recognition model is scheduled to recognize the enhanced real-time image data and determine the target recognition result in the real-time image data.
[0037] The enhanced real-time image data from step S103 is input into the underwater target recognition model trained in step S102. The model performs inference operations on the submersible's computing platform and quickly outputs the target recognition result for the current image. The recognition result typically includes: the detected target category and the target's location in the image (bounding box coordinates).
[0038] S105. Based on the target recognition results in the real-time image data, control the submersible to collect and store data, and complete the submersible data collection based on deep learning.
[0039] This application provides a deep learning-based data acquisition method for submersibles, integrating intelligent perception and decision-making control capabilities into the submersible system, thus enabling it to possess a high degree of autonomy. Even in deep-sea environments where communication is limited or impossible, the submersible can still independently complete complex target discovery, tracking, and refined data acquisition tasks. This not only reduces reliance on manual operation and remote communication, lowering operational costs and risks, but also significantly expands the submersible's operational range and duration, enabling it to undertake longer-duration, longer-distance, and more complex ocean exploration missions.
[0040] In one possible design approach, constructing a deep learning model includes: constructing a YOLO v3 model to obtain the deep learning model. It is worth noting that the YOLO v3 model is merely a preferred embodiment of this application; other models can also be used to construct the deep learning model.
[0041] In one possible design approach, the deep learning model is trained based on the historical image data and manually labeled tags corresponding to the historical image data to obtain an underwater target recognition model, including: The hyperparameters of the deep learning model are initialized and encoded to obtain the training population; The method for obtaining the training population can refer to existing technologies, and will not be described in detail in the embodiments of this application.
[0042] For any individual in the training population, the fitness of the individual is obtained by training based on the historical image data and manually labeled tags corresponding to the historical image data. Based on the fitness of the individual, determine the first optimal individual in the current training process; Based on the first optimal individual, a normal distribution local development strategy is used to locally develop the individual to obtain the locally developed individual; An information transmission strategy is used to perform information transmission development on the individuals after the partial development, and to obtain the individuals after the information transmission development; A normal distribution global development strategy is adopted to perform global development on the individuals after the information transmission development, and the individuals after global development are obtained. Determine if the current number of training iterations is greater than or equal to the preset maximum number of training iterations. If so, obtain the second optimal individual based on the individuals after global development; otherwise, return to the step of obtaining the first optimal individual. The hyperparameters of the second optimal individual are used as the final hyperparameters of the deep learning model to obtain the underwater target recognition model.
[0043] In one possible design approach, the fitness of an individual is obtained by training based on the historical image data and manually labeled tags corresponding to the historical image data, including: Using the historical image data as input data, and training with manually labeled tags corresponding to the historical image data as the expected output, the loss function value is obtained; The loss function value is added to a preset constant (such as 0.001 or 0.0001) to obtain a non-zero data item, and the reciprocal of the non-zero data item is taken to obtain the fitness of the individual.
[0044] In one possible design approach, based on the first optimal individual, a normal distribution local development strategy is used to locally develop the individual, obtaining the locally developed individual, including: The average position of the training population in the solution space is obtained as follows:
[0045] in, Indicates the first k The first training session i The first individual d The hyperparameter is N, which represents the total number of individuals in the training population. d =1,2,…,D, where D represents the total dimension of hyperparameters in an individual. Indicates the first k The average position during the training cycle. d dimensional hyperparameters; Based on the average position and the first optimal individual, obtain the first... i The simulated mean information for each individual is:
[0046] in, Indicates the first k During the training session i The simulated average information corresponding to each individual d dimensional hyperparameters, The first optimal individual represents the first optimal individual. d dimensional hyperparameters; Based on the average position, the first optimal individual, and the simulated mean information, obtain the first... i The simulated standard deviation information for each individual is as follows:
[0047] in, Indicates the first i The simulated standard deviation information for each individual. d dimensional hyperparameters; The penalty factor is:
[0048] in, Indicates the penalty factor. Represents the first random number between (0,1). This represents the second random number between (0,1). This represents a third random number between (0,1). This represents the fourth random number between (0,1). Represents the cosine function. Represents pi (π). Represents a logarithmic function; Based on the simulated mean information, simulated standard deviation information, and penalty factor, individuals are locally developed to obtain the following individuals after local development:
[0049] in, Indicates the first i The first individual after partial development d Dimensional hyperparameters.
[0050] By introducing the population average position, optimal individual information, and a dynamically changing penalty factor, individuals are guided to perform a refined normal distribution search near the current optimal solution. This strategy can quickly converge to the local optimum and explore it in depth, improving the accuracy and efficiency of optimization.
[0051] In one possible design approach, an information delivery strategy is employed to perform information delivery development on the individuals after the partial development, and the individuals after information delivery development are obtained, including: The information transmission vector is:
[0052]
[0053] in, Indicates the first j The information transmission vector corresponding to the individual after partial development is the first d dimensional hyperparameters, Indicates adaptive inertia weights, This represents the fifth random number that is uniformly distributed between (0,1). Represents the sixth random number uniformly distributed between (0,1). Indicates the first kThe first training session i After the individual's partial development, the first individual's d dimensional hyperparameters, Indicates the first j The Euclidean distance between an individual after partial development and the first optimal individual; Indicates the maximum number of training iterations. This represents a random number that follows a standard normal distribution and is distributed according to the first normal distribution. This represents the seventh random number between (0,1). This represents the eighth random number between (0, 1); Represents the first individual after random local development. d dimensional hyperparameters; Based on the information transmission vector, the information transmission individual is obtained as follows:
[0054] in, Indicates the first j The information transmission of the individual corresponding to the individual after partial development. d dimensional hyperparameters; Based on the information transmission individual, information transmission development is performed on the individual after the partial development, and the individual after information transmission development is obtained as follows:
[0055] in, Indicates the first j An individual after information transmission development, Indicates the first j The fitness of an individual after partial development. Indicates the first j The fitness of an individual after partial development, corresponding to information transmission. Indicates the first j The Euclidean distance between an individual after partial development and the individual transmitting information.
[0056] Individuals not only learn from the current best individual, but also exchange information with random individuals and adjust through adaptive inertia weights and distance factors. This effectively avoids the rapid loss of population diversity, prevents the algorithm from falling into "premature convergence" and local optima, and enhances the algorithm's ability to escape local optima.
[0057] In one possible design approach, a normal distribution global development strategy is employed to globally develop the individuals after the information transmission development, obtaining the globally developed individuals, including: For each individual that has undergone information transmission development, three other individuals that have undergone information transmission development are randomly matched to obtain the first tracking individual, the second tracking individual, and the third tracking individual corresponding to each individual that has undergone information transmission development. Based on the first tracking individual, the second tracking individual, and the third tracking individual, the first tracking vector and the second tracking vector are obtained as follows:
[0058]
[0059] in, Indicates the first k During the training session m The first individual after information transmission development d dimensional hyperparameters, This indicates the first individual being tracked. d dimensional hyperparameters, Represents the first tracking vector's... d dimensional hyperparameters, Indicates the first k During the training session m The fitness of an individual after the development of information transmission. Indicates the fitness of the first tracked individual; The second tracking vector represents the first... d dimensional hyperparameters, The second tracking vector represents the first... d dimensional hyperparameters, The third tracking vector represents the... d dimensional hyperparameters, This indicates the fitness of the second tracked individual. This indicates the fitness of the third tracked individual; Based on the first tracking vector and the second tracking vector, global development is performed on the individual after the information transmission development, and the individual after global development is obtained as follows:
[0060] in, Indicates the first m The first individual after global development d dimensional hyperparameters, Represents the tenth random number between (0,1). This represents a random number that follows a standard normal distribution or a second normal distribution. This represents a random number that follows a standard normal distribution but is distributed in the third normal distribution.
[0061] Optionally, a greedy strategy can be used to control the global development process, and after each strategy is executed, out-of-bounds handling can be performed on individuals to ensure that the hyperparameters always stay within their upper and lower limits.
[0062] By randomly selecting multiple individuals and constructing tracking vectors based on their fitness differences, the algorithm guides these individuals to explore a wide range of solutions. The introduction of a normal distribution increases the randomness and breadth of the search, ensuring that the algorithm can effectively explore globally and thus has a greater probability of finding the global optimum.
[0063] By organically combining and iteratively executing these three strategies, the embodiments of this application can automatically and efficiently find an optimal set of hyperparameters for the YOLOv3 model, so that the trained underwater target recognition model outperforms the model trained by traditional methods in key indicators such as recognition accuracy, recall, and mean precision, thereby maximizing model performance and ensuring the accuracy and completeness of data collection, effectively avoiding data omissions.
[0064] In one possible design approach, the real-time image data is subjected to image enhancement processing to obtain enhanced real-time image data, including: The real-time image data is subjected to three-channel color compensation, global histogram stretching, and / or HSV space enhancement to obtain enhanced real-time image data.
[0065] Three-channel color compensation: Based on the absorption and attenuation patterns of red, green, and blue light at different depths, the RGB channels of the image are compensated for with varying degrees of gain, especially enhancing the severely absorbed red light channel to restore the true colors of objects. For example, red-blue channel equalization can be performed on underwater images.
[0066] Global histogram stretching: Calculates the global histogram of the image and stretches it across the entire dynamic range to increase the overall contrast of the image, making dark details clearer and bright areas less overexposed. For example, processed three-channel values can be linearly mapped to the 0-1 range to complete the overall normalization process, thereby improving the visibility of low-light areas in the image and compensating for color attenuation, thus enhancing contrast.
[0067] HSV space enhancement processing: Converts the image from RGB space to HSV (hue, saturation, brightness) space. The saturation channel is moderately enhanced to make colors more vivid; gamma correction or histogram equalization is performed on the brightness channel to improve uneven lighting.
[0068] It is worth noting that the above three image enhancement methods are merely preferred embodiments of this application. Other image enhancement methods can also be used to process image data, thereby improving the accuracy of data recognition.
[0069] This application incorporates image enhancement processing in both the model training and real-time application phases. Before training, historical image data is enhanced to effectively eliminate color and brightness differences between images from different batches and environments, constructing a more unified and high-quality dataset. This ensures the model learns essential features rather than environmental noise. During real-time acquisition, the original images undergo the same three-channel color compensation, global histogram stretching, and HSV space enhancement processing. This effectively corrects color distortion in underwater images, enhances contrast and detail, and significantly improves the image quality input to the recognition model. This end-to-end image enhancement strategy fundamentally overcomes the interference of water optical characteristics on visual recognition, enabling the YOLO v3 model to maintain high accuracy and stability in target recognition within complex and variable underwater environments.
[0070] In one possible design approach, based on the target recognition results in the real-time image data, the submersible is controlled to perform data acquisition and storage, including: If the target recognition result in the real-time image data contains a preset target, the data sampling frequency is increased, and the increased data sampling frequency is used to collect real-time image data, or the recording function is enabled to collect real-time video data, and the collected real-time image data and real-time video data are stored.
[0071] This application establishes an intelligent closed loop of "identification-decision-action". The submersible no longer blindly executes preset programs, but uses its onboard high-performance underwater target recognition model to perceive and understand the environment in real time. When a preset target of interest (such as rare organisms, archaeological sites, minerals, etc.) is identified, the system can immediately make an intelligent decision.
[0072] Dynamically adjust sampling strategy: automatically increase image sampling frequency, switch from "slow scan" to "high speed gaze" to capture more instantaneous, high-resolution details of the target.
[0073] Trigger advanced acquisition mode: Automatically activates the recording function, upgrading discrete image data into a continuous video stream, fully recording the target's dynamic behavior or interaction with the environment, which has irreplaceable value for biological research, target dynamic analysis, etc.
[0074] This on-demand, intelligently triggered working mode has completely changed the traditional data acquisition method of submersibles, reducing redundant data acquisition in valueless areas and saving up to 70% more in storage space and energy consumption. At the same time, it ensures comprehensive and high-quality capture of key target data, significantly improving the scientific and commercial value density of the collected data.
[0075] like Figure 2 As shown, based on the same inventive concept, this application also provides a deep learning-based submersible data acquisition device, comprising: The training data acquisition module 201 is used to acquire historical image data collected by the submersible at historical moments and manually labeled tags corresponding to the historical image data; wherein, the historical image data is data after image enhancement processing; Deep learning model 202 is used to construct a deep learning model and train the deep learning model based on the historical image data and manually labeled labels corresponding to the historical image data to obtain an underwater target recognition model; The image enhancement module 203 is used to collect real-time image data at a preset data sampling frequency during the underwater operation of the submersible, and to perform image enhancement processing on the real-time image data to obtain enhanced real-time image data. The target recognition module 204 is used to schedule the underwater target recognition model to recognize the enhanced real-time image data and determine the target recognition result in the real-time image data; The submersible data acquisition module 205 is used to control the submersible to acquire and store data based on the target recognition results in the real-time image data, thereby completing the deep learning-based submersible data acquisition.
[0076] The data acquisition device for a submersible based on deep learning provided in this application embodiment can perform the above-described method and technical solution. Its principle and beneficial effects are similar, and will not be described again here.
[0077] This application also provides a computer-readable storage medium storing a computer program / instructions thereon, which, when executed by a processor, implements the steps in the method disclosed in this application.
[0078] This application also provides a computer program product that, when run on an electronic device, causes a processor to execute the steps in the method disclosed in this application.
[0079] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0080] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, electronic devices, and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0081] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0082] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0083] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.
[0084] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0085] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A deep learning-based submersible data acquisition method, characterized in that, include: Acquire historical image data collected by the submersible at historical moments and manually labeled tags corresponding to the historical image data; wherein, the historical image data is data after image enhancement processing; A deep learning model is constructed and trained based on the historical image data and manually labeled information corresponding to the historical image data to obtain an underwater target recognition model. During the underwater operation of the submersible, real-time image data is acquired using a preset data sampling frequency, and the real-time image data is subjected to image enhancement processing to obtain enhanced real-time image data. The underwater target recognition model is scheduled to recognize the enhanced real-time image data to determine the target recognition result in the real-time image data; Based on the target recognition results in the real-time image data, the submersible is controlled to collect and store data, thus completing the deep learning-based submersible data collection. The deep learning model is trained based on the historical image data and manually labeled tags corresponding to the historical image data to obtain an underwater target recognition model, including: The hyperparameters of the deep learning model are initialized and encoded to obtain the training population; For any individual in the training population, the fitness of the individual is obtained by training based on the historical image data and manually labeled tags corresponding to the historical image data. Based on the fitness of the individual, determine the first optimal individual in the current training process; Based on the first optimal individual, a normal distribution local development strategy is used to locally develop the individual to obtain the locally developed individual; An information transmission strategy is used to perform information transmission development on the individuals after the partial development, and to obtain the individuals after the information transmission development; A normal distribution global development strategy is adopted to perform global development on the individuals after the information transmission development, and the individuals after global development are obtained. Determine if the current number of training iterations is greater than or equal to the preset maximum number of training iterations. If so, obtain the second optimal individual based on the individuals after global development; otherwise, return to the step of obtaining the first optimal individual. The hyperparameters of the second optimal individual are used as the final hyperparameters of the deep learning model to obtain the underwater target recognition model.
2. The deep learning-based submersible data acquisition method according to claim 1, characterized in that, Building a deep learning model includes: building the YOLO v3 model to obtain the deep learning model.
3. The deep learning-based submersible data acquisition method according to claim 1, characterized in that, Based on the historical image data and manually labeled tags corresponding to the historical image data, the fitness of an individual is obtained, including: Using the historical image data as input data, and training with manually labeled tags corresponding to the historical image data as the expected output, the loss function value is obtained; The loss function value is added to a preset constant to obtain a non-zero data item, and the reciprocal of the non-zero data item is taken to obtain the fitness of the individual.
4. The deep learning-based submersible data acquisition method according to claim 1, characterized in that, Based on the first optimal individual, a normal distribution local development strategy is used to locally develop the individual, obtaining the locally developed individual, including: The average position of the training population in the solution space is obtained as follows: in, Indicates the first k The first training session i The first individual d The hyperparameter is N, which represents the total number of individuals in the training population. d =1,2,…,D, where D represents the total dimension of hyperparameters in an individual. Indicates the first k The average position during the training cycle. d dimensional hyperparameters; Based on the average position and the first optimal individual, obtain the first... i The simulated mean information for each individual is: in, Indicates the first k During the training session i The simulated average information corresponding to each individual. d dimensional hyperparameters, The first optimal individual represents the first optimal individual. d dimensional hyperparameters; Based on the average position, the first optimal individual, and the simulated mean information, obtain the first... i The simulated standard deviation information for each individual is as follows: in, Indicates the first i The simulated standard deviation information for each individual. d dimensional hyperparameters; The penalty factor is: in, Indicates the penalty factor. Represents the first random number between (0,1). This represents the second random number between (0,1). This represents a third random number between (0,1). This represents the fourth random number between (0,1). Represents the cosine function. Represents pi (π). Represents a logarithmic function; Based on the simulated mean information, simulated standard deviation information, and penalty factor, individuals are locally developed to obtain the following individuals after local development: in, Indicates the first i The first individual after partial development d Dimensional hyperparameters.
5. The deep learning-based submersible data acquisition method according to claim 4, characterized in that, An information delivery strategy is employed to perform information delivery development on the individuals after the partial development, and to obtain the individuals after the information delivery development, including: The information transmission vector is: in, Indicates the first j The information transmission vector corresponding to the individual after partial development is the first d dimensional hyperparameters, Indicates adaptive inertia weights, This represents the fifth random number that is uniformly distributed between (0,1). Represents the sixth random number uniformly distributed between (0,1). Indicates the first k The first training session i After the individual's partial development, the first individual's d dimensional hyperparameters, Indicates the first j The Euclidean distance between an individual after partial development and the first optimal individual; Indicates the maximum number of training iterations. This represents a random number that follows a standard normal distribution and is distributed according to the first normal distribution. This represents the seventh random number between (0,1). This represents the eighth random number between (0, 1); Represents the first individual after random local development. d dimensional hyperparameters; Based on the information transmission vector, the information transmission individual is obtained as follows: in, Indicates the first j The information transmission of the individual corresponding to the individual after partial development. d dimensional hyperparameters; Based on the information transmission individual, information transmission development is performed on the individual after the partial development, and the individual after information transmission development is obtained as follows: in, Indicates the first j An individual after information transmission development, Indicates the first j The fitness of an individual after partial development. Indicates the first j The fitness of an individual after partial development, corresponding to information transmission. Indicates the first j The Euclidean distance between an individual after partial development and the individual transmitting information.
6. The deep learning-based submersible data acquisition method according to claim 5, characterized in that, A normal distribution global development strategy is used to perform global development on the individuals after the information transmission development, and the individuals after global development are obtained, including: For each individual that has undergone information transmission development, three other individuals that have undergone information transmission development are randomly matched to obtain the first tracking individual, the second tracking individual, and the third tracking individual corresponding to each individual that has undergone information transmission development. Based on the first tracking individual, the second tracking individual, and the third tracking individual, the first tracking vector and the second tracking vector are obtained as follows: in, Indicates the first k During the training session m The first individual after information transmission development d dimensional hyperparameters, This indicates the first individual being tracked. d dimensional hyperparameters, Represents the first tracking vector's... d dimensional hyperparameters, Indicates the first k During the training session m The fitness of an individual after the development of information transmission. Indicates the fitness of the first tracked individual; The second tracking vector represents the first... d dimensional hyperparameters, The second tracking vector represents the first... d dimensional hyperparameters, The third tracking vector represents the... d dimensional hyperparameters, This indicates the fitness of the second tracked individual. This indicates the fitness of the third tracked individual; Based on the first tracking vector and the second tracking vector, global development is performed on the individual after the information transmission development, and the individual after global development is obtained as follows: in, Indicates the first m The first individual after global development d dimensional hyperparameters, Represents the tenth random number between (0,1). This represents a random number that follows a standard normal distribution or a second normal distribution. This represents a random number that follows a standard normal distribution but is distributed in the third normal distribution.
7. The deep learning-based submersible data acquisition method according to claim 1, characterized in that, The real-time image data is subjected to image enhancement processing to obtain enhanced real-time image data, including: The real-time image data is subjected to three-channel color compensation, global histogram stretching, and / or HSV space enhancement to obtain enhanced real-time image data.
8. The deep learning-based submersible data acquisition method according to claim 1, characterized in that, Based on the target recognition results in the real-time image data, the submersible is controlled to perform data acquisition and storage, including: If the target recognition result in the real-time image data contains a preset target, the data sampling frequency is increased, and the increased data sampling frequency is used to collect real-time image data, or the recording function is enabled to collect real-time video data, and the collected real-time image data and real-time video data are stored.
9. A deep learning-based submersible data acquisition device, wherein the deep learning-based submersible data acquisition device is capable of executing the deep learning-based submersible data acquisition method according to any one of claims 1 to 8, characterized in that, include: The training data acquisition module is used to acquire historical image data collected by the submersible at historical moments and manually labeled tags corresponding to the historical image data; wherein, the historical image data is data after image enhancement processing; A deep learning model is used to construct a deep learning model and train the deep learning model based on the historical image data and manually labeled tags corresponding to the historical image data to obtain an underwater target recognition model. The image enhancement module is used to acquire real-time image data at a preset data sampling frequency during the underwater operation of the submersible, and to perform image enhancement processing on the real-time image data to obtain enhanced real-time image data. The target recognition module is used to schedule the underwater target recognition model to recognize the enhanced real-time image data and determine the target recognition result in the real-time image data. The submersible data acquisition module is used to control the submersible to acquire and store data based on the target recognition results in the real-time image data, thereby completing the deep learning-based submersible data acquisition.
Citation Information
Patent Citations
Underwater image enhancement method and device based on self-supervision, and computer storage medium
CN115100063A
Underwater image enhancement processing method and system based on correlated imaging
CN118747713A
Underwater multi-unmanned-platform cooperative tracking method and system based on deep reinforcement learning
CN118779073A
Navigation method of intelligent navigation robot based on multi-sensor fusion
CN118936472A