Method for detecting feeding intensity and related device, equipment and storage medium
By using a feeding detection model in deep-sea aquaculture to detect and optimize the feeding audio of aquatic products, the accuracy and stability issues of computer vision methods under natural conditions were solved, achieving higher accuracy and stability in feeding intensity detection.
Patent Information
- Application Number
- CN202411236647.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-04
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2044-09-04
AI Technical Summary
In deep-sea aquaculture, existing computer vision-based methods for detecting feeding intensity are affected by natural conditions such as clear water, sufficient light, and stable background, resulting in insufficient detection accuracy and stability.
A feeding detection model was used to detect the feeding audio of aquatic products in the target water area. The first feeding intensity check was obtained by training the sample audio set, the sample audio set was updated, and the training model was optimized based on the check results to adapt to the growth stage of aquatic products.
It improves the accuracy and stability of feeding intensity detection, especially in deep-sea aquaculture scenarios, reduces the impact of natural conditions on detection, and ensures that the model adapts to changes in the growth stages of aquatic products.
Smart Images

Figure CN119323970B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to a feeding intensity detection method and related device, equipment and storage medium. BACKGROUND
[0002] In aquaculture, especially in deep-sea aquaculture, accurate feeding of bait is an important part that cannot be ignored. Among them, evaluating feeding intensity is a key point in accurate feeding of bait.
[0003] At present, the method based on computer vision is usually used to evaluate the feeding intensity. However, this method usually takes the natural conditions such as clear water, sufficient light and stable background as the premise, otherwise the detection accuracy of the feeding intensity will be greatly affected. Therefore, how to improve the accuracy and stability of detecting the feeding intensity, especially in the deep-sea aquaculture scene, has become a problem to be solved. SUMMARY
[0004] The technical problem solved by the present application is to provide a feeding intensity detection method and related device, equipment and storage medium, which can improve the accuracy and stability of detecting the feeding intensity, especially in the deep-sea aquaculture scene.
[0005] In order to solve the above technical problem, the first aspect of the present application provides a feeding intensity detection method, comprising: detecting the feeding audio of the target aquatic product in the target water area based on a feeding detection model to obtain a first feeding intensity; wherein the feeding detection model is obtained based on a sample audio set; obtaining a second feeding intensity after the first feeding intensity is checked; updating the sample audio set based on the feeding audio labeled with the second feeding intensity as a sample audio, obtaining a model analysis result representing whether the current growth stage of the target aquatic product is adapted to the feeding detection model based on the checking statistical result of the feeding intensity; determining whether to optimize and train the feeding detection model based on the sample audio set based on the model analysis result.
[0006] To solve the above technical problems, the second aspect of the present application provides a feeding intensity detection device, comprising: a feeding detection module, a strength checking module, an update analysis module and an optimization training module, the feeding detection module is used for detecting the feeding audio of the target aquatic product in the target water area based on the feeding detection model to obtain the first feeding intensity; wherein the feeding detection model is obtained based on the sample audio set; the strength checking module is used for obtaining the second feeding intensity after the first feeding intensity is checked; the update analysis module is used for updating the sample audio set based on the feeding audio labeled with the second feeding intensity and as a sample audio, obtaining the model analysis result representing whether the current growth stage of the target aquatic product is adapted to the feeding detection model based on the checking statistical result of the feeding intensity; the optimization training module is used for determining whether to optimize and train the feeding detection model based on the sample audio set based on the model analysis result.
[0007] To solve the above technical problems, the third aspect of the present application provides an electronic device, at least comprising a memory and a processor coupled with each other, the memory at least stores program instructions, and the processor is used to execute the program instructions to realize the feeding intensity detection method in the first aspect.
[0008] To solve the above technical problems, the fourth aspect of the present application provides a computer readable storage medium, which stores program instructions capable of being executed by a processor, and the program instructions are used to realize the feeding intensity detection method of the first aspect.
[0009] The above scheme detects the feeding audio of the target aquatic product in the target water area based on the feeding detection model to obtain the first feeding intensity, and the feeding detection model is obtained based on the sample audio set, then the second feeding intensity after the first feeding intensity is checked is obtained, to update the sample audio set based on the feeding audio labeled with the second feeding intensity and as a sample audio, and obtain the model analysis result representing whether the current growth stage of the target aquatic product is adapted to the feeding detection model based on the checking statistical result of the feeding intensity, and determine whether to optimize and train the feeding detection model based on the sample audio set based on the model analysis result, so on the one hand, it is distinguished from using computer vision technology and using audio data for feeding detection, which can as much as possible avoid the influence of natural conditions such as whether the water is clear, whether the light is sufficient, and whether the background is stable on the feeding detection, and only the feeding sound needs to be concerned, and the feeding sound usually exists stably in the feeding process of the aquatic product, on the other hand, the feeding detection model is updated and trained based on the statistical analysis of the feeding checking during the feeding detection, which can make the applicability of the feeding detection model as much as possible to follow the growth stage of the aquatic product. Therefore, the accuracy and stability of detecting the feeding intensity can be improved, especially in the deep sea aquaculture scene. BRIEF DESCRIPTION OF DRAWINGS
[0010] Figure 1 is a flowchart of an embodiment of the food intake intensity detection method of the present application;
[0011] Figure 2 is a framework diagram of an embodiment of the food intake detection model;
[0012] Figure 3 is a process diagram of an embodiment of the food intake intensity detection method of the present application;
[0013] Figure 4 is a framework diagram of an embodiment of the food intake intensity detection device of the present application;
[0014] Figure 5 is a framework diagram of an embodiment of the electronic device of the present application;
[0015] Figure 6 is a framework diagram of an embodiment of the computer-readable storage medium of the present application. DETAILED DESCRIPTION
[0016] The schemes of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0017] In the following description, specific details such as specific system structures, interfaces, techniques, etc. are presented in order to provide a thorough understanding of the present application for the sake of explanation, but not for the purpose of limiting the present application.
[0018] The terms "system" and "network" are often used interchangeably herein. The term "and / or" herein is merely an associative relationship between the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the three cases of A alone, A and B together, and B alone. In addition, the segment " / " herein generally represents an "or" relationship between the associated objects. In addition, "multiple" herein means two or more than two.
[0019] Please refer to Figure 1 , Figure 1 is a flowchart of an embodiment of the food intake intensity detection method of the present application. Specifically, it can include the following steps:
[0020] Step S11: detecting the food intake audio of the target aquatic product in the target water area based on the food intake detection model to obtain a first food intake intensity.
[0021] In the embodiments of the present disclosure, the feeding detection model is trained based on a sample audio set. It should be noted that the sample audio set can include feeding audio collected by the target aquatic product during feeding. For the sake of distinction, the feeding audio can also be referred to as a sample audio after being included in the sample audio set. The sample audio in the sample audio set can also be labeled with a feeding intensity. For the sake of distinction, the feeding intensity labeled by the sample audio can be referred to as a sample feeding intensity. In addition, as a possible example, the specific intensity of the “feeding intensity” such as “first feeding intensity”, “second feeding intensity”, “sample feeding intensity” and the like in the embodiments of the present disclosure can be any one of a plurality of preset feeding intensities. For example, the plurality of preset feeding intensities can include but are not limited to: no feeding, weak feeding, medium feeding, strong feeding, and the like. Of course, the above examples are only one possible example of a plurality of feeding intensities in the actual application, and do not limit the specific content of the plurality of feeding intensities. For example, the plurality of feeding intensities can also be divided more finely or more coarsely, which is not limited herein.
[0022] In one implementation scenario, the target aquatic product can include but is not limited to fish, shellfish, shrimp, crab, turtle, frog, etc., and the specific type of the target aquatic product is not limited herein.
[0023] In one implementation scenario, the acoustic features (such as Fbank, MFCC, etc.) of the feeding audio can be extracted, and then the acoustic features of the feeding audio are input into the feeding detection model, so that the first feeding intensity output by the feeding detection model can be obtained. It should be noted that the feeding detection model can include but is not limited to a convolutional neural network, and the network architecture of the feeding detection model is not limited herein. In order to improve the accuracy of the feeding detection model, the feeding detection model can also be pre-trained. As a possible example, a sample audio set can be obtained in advance, and the sample audio in the sample audio set is the feeding audio collected by the target aquatic product. The sample audio can be labeled with a sample feeding intensity. On this basis, the sample audio in the sample audio set can be detected based on the feeding detection model to obtain the predicted feeding intensity of the sample audio (for example, the acoustic features of the sample audio can be input into the feeding detection model to obtain the predicted feeding intensity output by the feeding detection model), so that the difference between the sample feeding intensity and the predicted feeding intensity can be measured by using a loss function such as cross-entropy, and the training loss of the feeding detection model can be obtained, and the network parameters of the feeding detection model can be adjusted based on the training loss. It should be noted that the specific measurement of the loss function can refer to the technical details of the loss function such as cross-entropy, which is not described herein.
[0024] In one specific implementation scenario, as one possible example, the target aquatic product feeding sound signal data can be collected by N hydrophones carried by the culture net cage platform within the time range of the target aquatic product feeding, amplified, filtered, and format-converted by the signal conditioner, and then stored as the feeding audio (in the training stage, which can be used as a sample feeding audio). In addition, in order to facilitate data labeling of the sample feeding audio, the feeding video signal can also be recorded synchronously as auxiliary information to form a sample audio set when conditions permit:
[0025] …(1)
[0026] In the above formula (1), represents the sample audio set, represents the i-th sample feeding audio in the sample audio set, represents the total number of sample feeding audios in the sample audio set. Of course, the above example is only one possible example of obtaining feeding audio, and does not limit the collection of feeding audio by other means, and will not be repeated here. It should be noted that the initially collected sample audio set can be regarded as basic data for initial training of the feeding detection model, and the updated sample audio set can be regarded as update data for optimizing and updating the feeding detection model. For details, please refer to the subsequent description, which will not be described here.
[0027] In one specific implementation scenario, in order to facilitate data labeling of the sample feeding audio, a knowledge base of the target aquatic product related to feeding intensity can also be constructed in advance. The knowledge base can define the definition information of the target aquatic product related to various feeding intensities. Taking the target aquatic product as fish and the feeding intensity as the four levels as described above as an example, the definition information of "strong feeding" can include but is not limited to: "fish school reacts strongly to food, gathers near the feeding point, and feeding behavior produces relatively large sound", the definition information of "medium feeding" can include but is not limited to: "fish school reacts strongly to food, has gathering behavior, and feeding behavior produces certain sound", the definition information of "weak feeding" can include but is not limited to: "weak reaction to food, only a small part of feeding behavior occurs, fish school sound is weak and slightly fluctuates", and the definition information of "no feeding" can include but is not limited to: "fish school has no reaction to food, and fish school sound is weak and smooth". Of course, the above knowledge base is only one possible example in the actual application process, and other cases can be inferred accordingly, and will not be repeated here.
[0028] In another implementation scenario, different from the implementation of the aforementioned feeding detection model, as another possible implementation of the feeding detection model, feature spectrum graphs of several feature categories of the feeding audio can also be extracted, and sub-weights of the various feature spectrum graphs in the same frequency band are obtained, and then the sub-spectrum graphs of the various feature spectrum graphs in the corresponding frequency band are weighted based on the sub-weights of the various feature spectrum graphs in the same frequency band to obtain weighted spectrum graphs, and the first feeding intensity is obtained based on the weighted spectrum graphs. The above-mentioned method can not only comprehensively utilize the advantages of various feature spectrum graphs in capturing feeding activity information, but also can perform spectrum graph fusion in the frequency band dimension, which helps to improve the spectrum graph fusion granularity, and further improves the detection accuracy of the feeding intensity.
[0029] In a specific implementation scenario, taking the feature spectrum graphs of several feature categories as CQT (Constant-Q Transform, constant-Q transform) spectrum graphs for example, the CQT spectrum graph can be represented as:
[0030] … (2)
[0031] In the above formula (2), is a window function with a length of Q is a constant factor in CQT, and k is the frequency sequence number of the CQT spectrum graph. For details of CQT, please refer to the technical details of CQT, which will not be repeated here.
[0032] In a specific implementation scenario, taking the feature spectrum graphs of several feature categories as STFT (Short Time Fourier Transform, short-time Fourier transform) spectrum graphs for example, the STFT spectrum graph can be represented as:
[0033] … (3)
[0034] In the above formula (3), is a window function. By adjusting the position of the window to intercept the signal on the time axis for Fourier transform, the purpose of short-time Fourier transform spectrum analysis is achieved. For details of STFT, please refer to the technical details of STFT, which will not be repeated here.
[0035] In a specific implementation scenario, taking the feature spectrum graphs of several feature categories as MFCC (Mel Frequency Cepstral Coefficient, Mel frequency cepstral coefficient) spectrum graphs for example, the specific extraction process includes the following steps: preprocessing, Fourier transform, Mel filter bank, logarithmic operation, discrete cosine transform, and dynamic feature extraction. Among them, the most critical is the design of the Mel filter bank, and each filter function can be represented as:
[0036] … (4)
[0037] In the above formula (4), is the center frequency of each filter, which can be referred to technical details of MFCC and will not be described here.
[0038] In a specific implementation scenario, taking the feature spectrograms of several feature categories including CQT spectrogram, STFT spectrogram and MFCC spectrogram as an example, the feature spectrograms of several feature categories of the feeding audio can be uniformly expressed as:
[0039] … (5)
[0040] In the above formula (5), denotes the feature spectrograms of several feature categories of the feeding audio, denotes the CQT spectrogram of the feeding audio, denotes the STFT spectrogram of the feeding audio, denotes the MFCC spectrogram of the feeding audio. In addition, the sample feeding audio with unclear feature expression and quality problems in the sample audio set can be cleaned.
[0041] In a specific implementation scenario, in order to measure the sub-weights of various feature spectrograms in the same frequency band, the energy intensity of various feature spectrograms in the same frequency band can be obtained, and each feature spectrogram can be selected as a current spectrogram, and the feature spectrogram other than the current spectrogram can be selected as a reference spectrogram, so that the sub-weight of the current spectrogram in the corresponding frequency band can be obtained based on the ratio of the energy intensity of the current spectrogram and the reference spectrogram in the same frequency band. It should be noted that in the case that there are multiple feature spectrograms other than the current spectrogram, the sub-weight of the current spectrogram in the corresponding frequency band can be obtained based on the ratio of the energy intensity of the current spectrogram in the specific frequency band and the sum of the energy intensities of each reference spectrogram in the corresponding frequency band. For ease of description, still taking the feature spectrograms of several feature categories including CQT spectrogram, STFT spectrogram and MFCC spectrogram as an example, the sub-weights of the CQT spectrogram, the STFT spectrogram and the MFCC spectrogram in the i-th frequency band can be respectively expressed as: In the above formula (6) to (8),
[0042] … (6)
[0043] … (7)
[0044] … (8)
[0045] In the above formula (6) to (8), , , respectively represent the energy intensity of the CQT spectrogram, the STFT spectrogram and the MFCC spectrogram respectively in the frequency band respectively represent the sub-weight of the CQT spectrogram, the STFT spectrogram and the MFCC spectrogram respectively in the frequency band
[0046] …(9)
[0047] In the above formula (9), represents the weight matrix, respectively represent the weight matrix of the CQT spectrogram, the STFT spectrogram and the MFCC spectrogram respectively in the weight matrix. It should be noted that the above example is only one possible example of obtaining the sub-weight of various feature spectrograms in the same frequency band, and does not limit other ways of obtaining the sub-weight. For example, the sub-weight of the current spectrogram in the corresponding frequency band can also be obtained based on the ratio of the energy intensity of the current spectrogram to the energy intensity of various feature spectrograms in the same frequency band, which will not be exemplified one by one here.
[0048] In one implementation scenario, as one possible example, please refer to Figure 2 , Figure 2 is a schematic diagram of the framework of an embodiment of the feeding detection model. As Figure 2 As shown, the feeding detection model can specifically adopt a design architecture of cascaded coupling of "feature extractor + feature fusioner + classifier". The implementation process of "feature extractor + feature fusioner" can refer to the foregoing related description, which will not be described here again. Of course, in the implementation process, it can also not be limited to this. For example, after extracting the feature spectrum of several feature categories, the feature spectrum of several feature categories can be input into a network model such as a convolutional neural network for feature extraction, and then the feature fusion is performed in the feature extraction process according to the foregoing related description of feature fusion, and the fused weighted spectrum is further processed by the subsequent network layer until the classification layer, that is, the classification layer can divide the feature mapping range according to the learned mapping relationship between the features and the feeding intensity, thereby classifying the specific feeding audio as an intensity classification. As a possible example, the feeding intensity evaluation classification can be implemented based on a ResNet-34 network, and the feature data after feature fusion is used as the input of the ResNet-34. After data dimension adjustment, it is first input into the 7*7 convolutional layer in the ResNet-34 network, and then the batch normalization layer is connected, and then the 3*3 maximum pooling downsampling layer is connected. Then, 4 modules composed of residual blocks can be used. Among them, Conv2_x adopts 3 residual layers, Conv3_x adopts 4 residual layers, Conv4_x adopts 6 residual layers, and Conv5_x adopts 3 residual layers. Finally, the output channel number is changed to 4 through the full connection layer, that is, the probability values of various feeding intensities such as "strong feeding", "medium feeding", "weak feeding" and "no feeding" can be output, so as to realize the evaluation classification of feeding intensity. It should be noted that the above example is only one possible architecture of the feeding detection model, and the specific structure of the feeding detection model is not limited here.
[0049] Step S12: obtaining the second feeding intensity after the first feeding intensity is checked.
[0050] In one implementation scenario, the feeding detection model can be deployed on the application side. After obtaining the first feeding intensity of the feeding audio, the application side can upload the feeding audio and its first feeding intensity to the upper management system together, so that the upper management system checks the first feeding intensity of the feeding audio. Get the second feeding intensity of the feeding audio. In addition, in the case of collecting feeding audio and also collecting feeding video, the feeding video can also be uploaded to the upper management system together with the feeding audio and its first feeding intensity, so that the upper management system checks the first feeding intensity of the feeding audio by referring to the feeding video. In addition, after the upper management system checks the first feeding intensity, the second feeding intensity of the feeding audio after the check can be fed back to the application side. As a possible example, the upper management system can be manned by relevant personnel with aquaculture professional knowledge, so that the relevant personnel check the first feeding intensity.
[0051] In one implementation scenario, the second feeding intensity after the verification can be the same as the first feeding intensity, or the second feeding intensity after the verification can also be different from the first feeding intensity, which is not limited herein.
[0052] Step S13: updating the sample audio set based on the feeding audio labeled with the second feeding intensity and as the sample audio, and obtaining the model analysis result characterizing whether the current growth stage of the target aquatic product is adapted to the feeding detection model based on the verification statistical result of the feeding intensity.
[0053] In one implementation scenario, after obtaining the second feeding intensity of the feeding audio, the second feeding intensity can be labeled on the feeding audio, and the feeding audio labeled with the second feeding intensity can be used as the sample audio to update the sample audio set.
[0054] In one specific implementation scenario, in order to update the sample audio set, the sample audio before the last optimization training of the feeding detection model in the sample audio set can be deleted. For example, as described above, the sample audio set used for the initial training of the feeding detection model (for the sake of distinction, it can be referred to as the basic data) can include the sample audio to After the initial training, the feeding detection model can be used to continue detecting the feeding intensity and the foregoing related steps on the newly collected feeding audio until the current growth stage of the target aquatic product is found to be not adapted to the feeding detection model for the first time, at which time the sample audio to before the first optimization training of the feeding detection model in the sample audio set, i.e., the sample audio to may be deleted. That is, the updated sample audio set (for the sake of distinction, it can be referred to as the updated data) can include the sample audio to . Of course, the above example is only one possible example, and other cases can be similarly extended, which will not be repeated herein. For example, only part of the sample audio to may be deleted, but not all. Specifically, the sample audio before the last optimization training of the feeding detection model in the sample audio set can be sorted in the order from the earliest to the latest, and the sample audio before the preset sequence (e.g., the first 50%, the first 70%, etc.; or the first 100%, i.e., all) can be deleted. The above method of deleting at least part of the sample audio before the last optimization training of the feeding detection model in the sample audio set can as much as possible eliminate the sample audio in the sample audio set that is not adapted to the current growth stage, which helps to improve the adaptation of the feeding detection model to the current growth stage after the optimization training.
[0055] In one specific implementation scenario, different from the foregoing implementation manners, in order to update the sample audio set, the sample audio before the last optimization training of the feeding detection model in the sample audio set can also be retained. For example, as described above, the sample audio set (for the sake of distinction, which can be referred to as basic data) used for initial training of the feeding detection model can contain sample audio to After initial training, the feeding detection model can be used to continue detecting feeding intensity and the foregoing related steps on newly collected feeding audio until the current growth stage of the target aquatic product is found to be not suitable for the feeding detection model, at which time the sample audio to has been newly accumulated. Then, the sample audio before the last optimization training of the feeding detection model in the sample audio set can be retained. That is, the sample audio to after updating the sample audio set contains sample audio to . As one possible example, in this case, when the feeding detection model is optimized and trained using the updated sample audio set, the learning rate of the sample audio to may be set to be lower than the learning rate of the sample audio to . In the above manner, the sample audio before the last optimization training of the feeding detection model in the sample audio set is retained, which can help the feeding detection model to support feeding detection at the growth stage level while referring to feeding audio up to the current growth stage, thereby helping the feeding detection model to more comprehensively support feeding detection at the growth stage level.
[0056] It should be noted that the above examples are only two possible examples of updating the sample audio set. In actual application, the sample audio set can be updated according to specific conditions, and is not limited to the above two examples. In addition, how to update the sample audio set can also be adaptively selected according to actual conditions. For example, in the case where the feeding performance of the target aquatic product changes greatly with the growth stage, the former of the above two ways can be selected, or in the case where the feeding performance of the target aquatic product changes relatively little with the growth stage, the latter of the above two ways can be selected.
[0057] In one implementation scenario, in order to obtain the model analysis result, each feeding audio after the last optimization training of the self-feeding detection model can be selected as a historical audio, and based on the checking result of each historical audio on the feeding intensity, a checking statistical result is obtained, and the checking result includes: a first audio in which the historical audio is consistent with the first feeding intensity and the second feeding intensity, or a second audio in which the historical audio is inconsistent with the first feeding intensity and the second feeding intensity, and the checking statistical result includes at least one of the following: a first proportion of the first audio, and a second proportion of the second audio. On this basis, the model analysis result can be obtained based on the checking statistical result. Through the above-mentioned manner, the checking statistical result of each historical audio on the feeding intensity after the last optimization training of the self-feeding detection model is obtained, and the model can be evaluated and analyzed after each checking according to the historical data, which helps to discover whether the model is suitable for the growth stage in a timely manner.
[0058] In one specific implementation scenario, as mentioned above, the sample audio set (for the sake of distinction, it can be referred to as basic data) used for initial training of the feeding detection model can include sample audios to After initial training, the feeding detection model can be used to continue detecting the feeding intensity and the foregoing related steps on the newly collected feeding audios, and in this process, the feeding audios , , and the like can be accumulated. For example, in the case of the currently collected feeding audios , the feeding audios to can be used as historical videos, respectively. Of course, the above example is only one possible example in the actual application process, and other cases can be similarly extended, which will not be exemplified one by one here.
[0059] In one specific implementation scenario, still taking the foregoing example as an example, the checking results of the feeding audios to can be obtained, that is, whether the first feeding intensity predicted by the feeding detection model is consistent with the second feeding intensity after checking. On this basis, the first proportion of the feeding audios (i.e., the first audio) in which the first feeding intensity is consistent with the second feeding intensity in the feeding audios to can be counted, and / or the second proportion of the feeding audios (i.e., the second audio) in which the first feeding intensity is inconsistent with the second feeding intensity in the feeding audios to can be counted. Of course, the above example is only one possible example in the actual application process, and other cases can be similarly extended, which will not be exemplified one by one here.
[0060] In a specific implementation scenario, in the case that the checking statistical result includes the first proportion, the model analysis result can be obtained based on a comparison result between the first proportion and the first threshold. Illustratively, in response to the first proportion being not lower than the first threshold, it can be determined that the model analysis result includes that the current growth stage of the target aquatic product is adapted to the feeding detection model. Conversely, in response to the first proportion being lower than the first threshold, it can be determined that the model analysis result includes that the current growth stage of the target aquatic product is not adapted to the feeding detection model. It should be noted that the specific value of the first threshold can be set according to the detection accuracy of the feeding intensity. Illustratively, in the case that the detection accuracy of the feeding intensity is relatively high, the first threshold can be set to be relatively large, or in the case that the detection accuracy of the feeding intensity is relatively loose, the first threshold can be set to be slightly smaller. The specific value of the first threshold is not limited herein. In the above manner, in the case that the checking statistical result includes the first proportion, the model analysis result is determined by obtaining the comparison result between the first proportion and the first threshold, which can determine the model analysis result by numerical comparison, and helps to simplify the evaluation complexity of the feeding detection model.
[0061] In a specific implementation scenario, in the case that the checking statistical result includes the second proportion, the model analysis result can be obtained based on a comparison result between the second proportion and the second threshold. Illustratively, in response to the second proportion being not lower than the second threshold, it can be determined that the model analysis result includes that the current growth stage of the target aquatic product is not adapted to the feeding detection model. Conversely, in response to the second proportion being lower than the second threshold, it can be determined that the model analysis result includes that the current growth stage of the target aquatic product is adapted to the feeding detection model. It should be noted that the specific value of the second threshold can be set according to the detection accuracy of the feeding intensity. Illustratively, in the case that the detection accuracy of the feeding intensity is relatively high, the second threshold can be set to be relatively small, or in the case that the detection accuracy of the feeding intensity is relatively loose, the second threshold can be set to be slightly larger. The specific value of the second threshold is not limited herein. In the above manner, in the case that the checking statistical result includes the second proportion, the model analysis result is determined by obtaining the comparison result between the second proportion and the second threshold, which can determine the model analysis result by numerical comparison, and helps to simplify the evaluation complexity of the feeding detection model.
[0062] Step S14: determining whether to perform optimized training on the feeding detection model based on the sample audio set based on the model analysis result.
[0063] In one implementation scenario, in response to the model analysis result indicating that the current growth stage is not adapted to the feeding detection model, the feeding detection model can be optimized and trained based on the sample audio set. The specific training process of the feeding detection model can be referred to the foregoing related description, which will not be described herein.Figure 3 , Figure 3 is a process schematic diagram of an embodiment of the present application. As shown in the figure, after the optimization training of the feeding detection model, the step of detecting the feeding audio of the target aquatic product in the target water area based on the feeding detection model can be returned to, to obtain the first feeding intensity. Such a loop iteration can follow the entire growth cycle of the target aquatic product and simultaneously optimize the training of the feeding detection model, which helps to achieve adaptive feeding detection during the entire growth cycle of the target aquatic product, so as to maintain the stability of the feeding intensity detection during the entire growth cycle. Figure 3
[0064] In one implementation scenario, in response to the model analysis result indicating that the current growth stage is suitable for the feeding detection model, the step of detecting the feeding audio of the target aquatic product in the target water area based on the feeding detection model can be returned to, so that the current feeding detection model can be used for feeding detection in the current growth stage, which helps to maintain the stability of the feeding intensity detection in the feeding detection result of the current growth stage.
[0065] It should be noted that the above loop iteration can continue until the target aquatic product in the target water area is caught, at which time there is no need to detect the feeding intensity any more; or, when the target aquatic product of the same species is re-introduced into the target water area after being caught, the feeding detection model can be initially trained using the basic data and deployed on the application side first, and then whether to optimize the training of the feeding detection model can be determined according to the checking statistical result of the feeding audio on the feeding intensity according to the foregoing related description; or, when a new target aquatic product (different from the original species) is re-introduced into the current target water area after the target aquatic product is caught, the feeding detection model can be initially trained using the basic data of the new target aquatic product and deployed on the application side first, and then whether to optimize the training of the feeding detection model can be determined according to the checking statistical result of the feeding audio on the feeding intensity according to the foregoing related description. The above examples are only a few possible examples in the actual application process, and do not limit other possible cases, which will not be exemplified one by one here.
[0066] The above scheme is based on the feeding detection model to detect the feeding audio of the target aquatic product in the target water area to obtain the first feeding intensity, and the feeding detection model is trained based on the sample audio set, and then the second feeding intensity after the first feeding intensity is checked is obtained, so as to update the sample audio set based on the feeding audio labeled with the second feeding intensity and as the sample audio, obtain the model analysis result representing whether the current growth stage of the target aquatic product is suitable for the feeding detection model based on the checking statistical result of the feeding intensity, and determine whether to optimize and train the feeding detection model based on the sample audio set based on the model analysis result. Therefore, on the one hand, the feeding detection is performed by using audio data instead of computer vision technology, which can avoid the influence of natural conditions such as whether the water is clear, whether the light is sufficient, and whether the background is stable on the feeding detection as much as possible, and only the feeding sound needs to be concerned. The feeding sound usually exists stably in the feeding process of the aquatic product. On the other hand, the feeding detection model is updated and trained based on the feeding checking statistical analysis, so that the applicability of the feeding detection model can follow the growth stage of the aquatic product as much as possible. Therefore, the accuracy and stability of detecting the feeding intensity can be improved, especially in the deep-sea aquaculture scene.
[0067] Please refer to Figure 4 , Figure 4 is a frame schematic diagram of an embodiment of the feeding intensity detection device. The feeding intensity detection device 40 comprises a feeding detection module 41, an intensity checking module 42, an updating analysis module 43, and an optimization training module 44. The feeding detection module 41 is used to detect the feeding audio of the target aquatic product in the target water area based on the feeding detection model to obtain the first feeding intensity. The feeding detection model is trained based on the sample audio set. The intensity checking module 42 is used to obtain the second feeding intensity after the first feeding intensity is checked. The updating analysis module 43 is used to update the sample audio set based on the feeding audio labeled with the second feeding intensity and as the sample audio, and obtain the model analysis result representing whether the current growth stage of the target aquatic product is suitable for the feeding detection model based on the checking statistical result of the feeding intensity. The optimization training module 44 is used to determine whether to optimize and train the feeding detection model based on the sample audio set based on the model analysis result.
[0068] The above scheme, the feeding intensity detection device 40 detects the feeding audio of the target aquatic product in the target water area based on the feeding detection model to obtain the first feeding intensity, and the feeding detection model is obtained based on the sample audio set training, and then the second feeding intensity after the first feeding intensity is checked is obtained, so as to update the sample audio set based on the feeding audio labeled with the second feeding intensity and as the sample audio, and obtain the model analysis result representing whether the current growth stage of the target aquatic product is suitable for the feeding detection model based on the checking statistical result of the feeding intensity, and determine whether to optimize and train the feeding detection model based on the sample audio set based on the model analysis result, so on the one hand, the feeding detection is performed by using audio data instead of computer vision technology, which can as much as possible avoid the influence of natural conditions such as whether the water is clear, whether the light is sufficient, and whether the background is stable on the feeding detection, and only the feeding sound needs to be concerned, and the feeding sound usually exists stably in the feeding process of the aquatic product, and on the other hand, the feeding detection model is updated and trained based on the feeding detection process and the statistical analysis of the feeding checking, and whether to update and train the feeding detection model is determined, so that the applicability of the feeding detection model can as much as possible follow the growth stage of the aquatic product. Therefore, the accuracy and stability of detecting the feeding intensity can be improved, especially in the deep-sea aquaculture scene.
[0069] In some disclosed embodiments, the optimization training module 44 includes an optimization training submodule for optimizing and training the feeding detection model based on the sample audio set in response to the model analysis result representing that the current growth stage is not suitable for the feeding detection model.
[0070] In some disclosed embodiments, the optimization training module 44 includes a first return submodule for returning to the step of detecting the feeding audio of the target aquatic product in the target water area based on the feeding detection model to obtain the first feeding intensity after performing the step of optimizing and training the feeding detection model based on the sample audio set.
[0071] In some disclosed embodiments, the optimization training module 44 includes a second return submodule for returning to the step of detecting the feeding audio of the target aquatic product in the target water area based on the feeding detection model to obtain the first feeding intensity in response to the model analysis result representing that the current growth stage is suitable for the feeding detection model.
[0072] In some disclosed embodiments, the update analysis module 43 comprises an audio selection sub-module configured to select each feeding audio since the last optimization training of the feeding detection model as a historical audio; the update analysis module 43 comprises a verification statistics sub-module configured to obtain verification statistics results based on verification results of each historical audio on feeding intensity; wherein the verification results comprise: the historical audio being a first audio with the first feeding intensity consistent with the second feeding intensity, or the historical audio being a second audio with the first feeding intensity inconsistent with the second feeding intensity, and the verification statistics results comprise at least one of: a first proportion of the first audio, and a second proportion of the second audio; and the update analysis module 43 comprises a result analysis sub-module configured to obtain model analysis results based on the verification statistics results.
[0073] In some disclosed embodiments, the result analysis sub-module comprises a first obtaining unit configured to, in a case where the verification statistics results comprise the first proportion, obtain the model analysis results based on a size comparison result between the first proportion and a first threshold; and the result analysis sub-module comprises a second obtaining unit configured to, in a case where the verification statistics results comprise the second proportion, obtain the model analysis results based on a size comparison result between the second proportion and a second threshold.
[0074] In some disclosed embodiments, the update analysis module 43 comprises a sample deletion sub-module configured to delete at least part of the sample audios in the sample audio set before the last optimization training of the feeding detection model; and the update analysis module 43 comprises a sample retention sub-module configured to retain the sample audios in the sample audio set before the last optimization training of the feeding detection model.
[0075] In some disclosed embodiments, the feeding detection module 41 comprises a feature extraction sub-module configured to extract feature spectrograms of a plurality of feature categories of the feeding audio; the feeding detection module 41 comprises a weight obtaining sub-module configured to obtain sub-weights of the feature spectrograms of the same frequency band; the feeding detection module 41 comprises a spectrogram weighting sub-module configured to weight the sub-spectrogram of the corresponding frequency band of each feature spectrogram based on the sub-weights of the feature spectrograms of the same frequency band, to obtain a weighted spectrogram; and the feeding detection module 41 comprises an intensity prediction sub-module configured to predict based on the weighted spectrogram to obtain the first feeding intensity.
[0076] In some disclosed embodiments, the weight obtaining sub-module comprises an energy intensity obtaining unit configured to obtain energy intensities of the feature spectrograms of the same frequency band; the weight obtaining sub-module comprises a feature spectrogram selection unit configured to select each feature spectrogram as a current spectrogram, and select a feature spectrogram other than the current spectrogram as a reference spectrogram; and the weight obtaining sub-module comprises a frequency band weight calculation unit configured to obtain the sub-weights of the current spectrogram in the corresponding frequency band based on a ratio of the energy intensities of the current spectrogram and the reference spectrogram in the same frequency band.
[0077] Please refer toFigure 5 , Figure 5 is a schematic diagram of a framework of an embodiment of the electronic device. The electronic device 50 at least includes a memory 51 and a processor 52 coupled to each other, the memory 51 at least stores program instructions, and the processor 52 is configured to execute the program instructions to implement the steps in any of the above-mentioned food intake intensity detection method embodiments. For details, refer to the foregoing embodiments, which will not be repeated here. As a possible example, the electronic device 50 can include but is not limited to a server, etc., and the specific type of the electronic device 50 is not limited here.
[0078] Specifically, the processor 52 is configured to control itself and the memory 51 to implement the steps in any of the above-mentioned food intake intensity detection method embodiments. The processor 52 can also be referred to as a CPU (Central Processing Unit). The processor 52 can be an integrated circuit chip with processing capability. The processor 52 can also be a general purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general purpose processor can be a microprocessor or the processor can also be any conventional processor. In addition, the processor 52 can be jointly implemented by an integrated circuit chip.
[0079] The above scheme is that the electronic device 50 detects the feeding audio of the target aquatic product in the target water area based on the feeding detection model to obtain the first feeding intensity, the feeding detection model is obtained based on the sample audio set, the second feeding intensity after the first feeding intensity is checked is obtained, the sample audio set is updated based on the feeding audio labeled with the second feeding intensity and as the sample audio, the model analysis result indicating whether the current growth stage of the target aquatic product is suitable for the feeding detection model is obtained based on the checking statistical result of the feeding intensity, and whether the feeding detection model is optimized and trained based on the sample audio set is determined based on the model analysis result. Therefore, on the one hand, the feeding detection is performed by using the audio data instead of the computer vision technology, so that the influence of natural conditions such as whether the water is clear, whether the light is sufficient, and whether the background is stable on the feeding detection can be avoided as much as possible, and only the feeding sound needs to be concerned. The feeding sound usually exists stably in the feeding process of the aquatic product. On the other hand, the feeding detection model is updated and trained based on the checking statistical result of the feeding in the feeding detection process, so that the applicability of the feeding detection model can follow the growth stage of the aquatic product as much as possible. Therefore, the accuracy and stability of the detection of the feeding intensity can be improved, especially in the deep-sea aquaculture scene.
[0080] Please refer to Figure 6 , Figure 6 is a framework schematic diagram of an embodiment of the computer readable storage medium 60 of the present application. The computer readable storage medium 60 stores program instructions 61 capable of being executed by a processor, and the program instructions 61 are used to implement the steps in any of the above feeding intensity detection method embodiments.
[0081] The above scheme, the computer readable storage medium 60 detects the feeding audio of the target aquatic product in the target water area based on the feeding detection model to obtain the first feeding intensity, and the feeding detection model is obtained based on the sample audio set training, and then the second feeding intensity after the first feeding intensity is checked is obtained, so as to update the sample audio set based on the feeding audio labeled with the second feeding intensity and as a sample audio, and based on the checking statistical result of the feeding intensity, the model analysis result representing whether the current growth stage of the target aquatic product is adapted to the feeding detection model is obtained, and based on the model analysis result, it is determined whether to optimize and train the feeding detection model based on the sample audio set, so on the one hand, the feeding detection is performed by using audio data instead of computer vision technology, which can as much as possible avoid the influence of natural conditions such as whether the water is clear, whether the light is sufficient, and whether the background is stable on the feeding detection, and only the feeding sound needs to be concerned. In general, the feeding sound is stable during the feeding process of the aquatic product, on the other hand, the feeding detection model is updated and trained based on the statistical analysis of the feeding checking during the feeding detection, which can make the applicability of the feeding detection model as much as possible to follow the growth stage of the aquatic product. Therefore, the accuracy and stability of detecting the feeding intensity can be improved, especially in the deep-sea aquaculture scene.
[0082] In some embodiments, the device provided by the embodiments of the present disclosure has functions or includes modules that can be used to perform the methods described in the above method embodiment descriptions, and the specific implementations can refer to the descriptions of the above method embodiments. For brevity, they will not be repeated here.
[0083] The above description of various embodiments tends to emphasize the differences between various embodiments, and the same or similar parts can be mutually referred to. For brevity, they will not be repeated here.
[0084] In several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only schematic, for example, the division of modules or units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed each other can be indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.
[0085] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, i.e., may be located in one place, or may be distributed to multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment.
[0086] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0087] If the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor (processor) to execute all or part of the steps of the methods of the various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0088] If the technical solutions of the present application involve personal information, the product applying the technical solutions of the present application has been explicitly informed of the personal information processing rules before processing the personal information, and has obtained the personal independent consent. If the technical solutions of the present application involve sensitive personal information, the product applying the technical solutions of the present application has obtained the personal independent consent before processing the sensitive personal information, and at the same time meets the requirement of "explicit consent". For example, at the personal information collection device such as camera, a clear and prominent mark is set to inform that the personal information collection range has been entered, and the personal information will be collected. If the person voluntarily enters the collection range, it is regarded as consent to collect the personal information; or on the device for processing personal information, the personal information processing rules are informed by using obvious marks / information, and the personal authorization is obtained by means of pop-up information or asking the person to upload his / her personal information, etc. The personal information processing rules can include personal information processor, personal information processing purpose, processing method, and personal information type, etc.
Claims
1. A method for detecting feeding intensity, characterized by, The method comprises the following steps: detecting feeding audio of a target aquatic product in a target water area based on a feeding detection model to obtain a first feeding intensity; wherein the feeding detection model is obtained based on a sample audio set; obtaining a second feeding intensity after the first feeding intensity is checked; updating the sample audio set based on feeding audio labeled with the second feeding intensity and serving as a sample audio, and obtaining a model analysis result representing whether a current growth stage of the target aquatic product is suitable for the feeding detection model based on a checking statistical result of the feeding intensity; determining whether to perform optimization training on the feeding detection model based on the sample audio set based on the model analysis result.
2. The method of claim 1, wherein, The step of determining whether to perform optimization training on the feeding detection model based on the sample audio set based on the model analysis result at least comprises: in response to the model analysis result representing that the current growth stage is not suitable for the feeding detection model, performing optimization training on the feeding detection model based on the sample audio set.
3. The method according to claim 1 or 2, characterized in that, After the step of performing optimization training on the feeding detection model based on the sample audio set, the method further comprises: returning to the step of detecting feeding audio of a target aquatic product in a target water area based on a feeding detection model to obtain a first feeding intensity.
4. The method of claim 1, wherein, The step of determining whether to perform optimization training on the feeding detection model based on the sample audio set based on the model analysis result at least comprises: in response to the model analysis result representing that the current growth stage is suitable for the feeding detection model, returning to the step of detecting feeding audio of a target aquatic product in a target water area based on a feeding detection model to obtain a first feeding intensity.
5. The method of claim 1, wherein, The step of obtaining a model analysis result representing whether a current growth stage of the target aquatic product is suitable for the feeding detection model based on a checking statistical result of the feeding intensity comprises: selecting each of the feeding audio since the last optimization training of the feeding detection model as a historical audio; obtaining the checking statistical result based on a checking result of each of the historical audio on the feeding intensity; wherein the checking result comprises: the historical audio being first audio in which the first feeding intensity is consistent with the second feeding intensity, or the historical audio being second audio in which the first feeding intensity is inconsistent with the second feeding intensity, and the checking statistical result comprises at least one of: a first proportion of the first audio, a second proportion of the second audio; obtaining the model analysis result based on the checking statistical result.
6. The method of claim 5, wherein, The step of obtaining the model analysis result based on the checking statistical result comprises at least one of: in a case where the checking statistical result comprises the first proportion, obtaining the model analysis result based on a comparison result between the first proportion and a first threshold value; in a case where the checking statistical result comprises the second proportion, obtaining the model analysis result based on a comparison result between the second proportion and a second threshold value.
7. The method of claim 1, wherein, The step of updating the sample audio set based on feeding audio labeled with the second feeding intensity and serving as a sample audio comprises any one of: delete at least part of the sample audio in the sample audio set before the last optimization training of the feeding detection model; retain the sample audio in the sample audio set before the last optimization training of the feeding detection model.
8. The method of claim 1, wherein, The feeding audio of the target aquatic product in the target water area is detected based on the feeding detection model to obtain a first feeding intensity, which includes: extracting feature spectrograms of several feature categories of the feeding audio; obtaining sub-weights of various feature spectrograms in the same frequency band; weighting sub-spectrograms of various feature spectrograms in the corresponding frequency band based on the sub-weights of various feature spectrograms in the same frequency band to obtain weighted spectrograms; based on the weighted spectrograms, the first feeding intensity is obtained.
9. The method of claim 8, wherein, The sub-weights of various feature spectrograms in the same frequency band are obtained, which includes: obtaining the energy intensity of various feature spectrograms in the same frequency band; selecting various feature spectrograms as current spectrograms respectively, and selecting the feature spectrograms other than the current spectrograms as reference spectrograms; based on the ratio of the energy intensity of the current spectrogram and the reference spectrogram in the same frequency band, the sub-weight of the current spectrogram in the corresponding frequency band is obtained.
10. An eating intensity detecting apparatus characterized by comprising: It includes: The feeding detection module is used for detecting the feeding audio of the target aquatic product in the target water area based on the feeding detection model to obtain a first feeding intensity; wherein the feeding detection model is obtained based on the sample audio set training; The intensity checking module is used for obtaining the second feeding intensity after the first feeding intensity is checked; The update analysis module is used for updating the sample audio set based on the feeding audio labeled with the second feeding intensity and as a sample audio, and obtaining a model analysis result representing whether the current growth stage of the target aquatic product is adapted to the feeding detection model based on the checking statistical result of the feeding intensity; The optimization training module is used for determining whether to perform optimization training on the feeding detection model based on the sample audio set based on the model analysis result.
11. An electronic device, comprising: At least including a memory and a processor coupled to each other, the memory at least stores program instructions, and the processor is used to execute the program instructions to realize the feeding intensity detection method of any one of claims 1-9.
12. A computer-readable storage medium, characterized in that, The program instructions capable of being run by the processor are stored, and the program instructions are used to realize the feeding intensity detection method of any one of claims 1-9.
Citation Information
Patent Citations
Broiler feed intake detection system based on audio technology
CN112331231A
Fish feeding behavior identification method based on YOLOv5
CN113537106A