Pet sound monitoring method

By encoding and multi-channel feature extraction of pet sounds, combined with convolutional neural networks and sequence processing methods, the problem of insufficient recognition accuracy in pet sound monitoring is solved, and the rapid and accurate identification of pet health status is achieved, and the flexibility and generalization ability of the model are improved.

CN120260587APending Publication Date: 2025-07-04VIENTIANE OPTICAL BIOTECHNOLOGY (CHENGDU) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510380189.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing pet sound monitoring technology lacks recognition accuracy when processing health warning signals such as pet breathing and screaming, and the existing methods fail to fully consider factors such as the frequency changes, duration and background noise of the sound, resulting in insufficient generalization capabilities of the model and it is difficult to maintain signal noise reduction robustness in complex sound fields.

Method used

By encoding pet sounds, three encoding units: α, β, and γ are generated. Random sampling and multi-channel feature extraction are used to construct training data sets, combined with convolutional neural networks for model training, and sequence processing methods are introduced to perform post-processing to improve recognition accuracy and interpretability.

Benefits of technology

Under a small amount of data labeling, the rapid and accurate identification of pet health status is achieved, the flexibility and generalization of the model are improved, and the vocal characteristics of pets can be identified in different health statuses, providing an efficient solution for pet health monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260587A_ABST
    Figure CN120260587A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of sound monitoring, and particularly relates to a pet sound monitoring method, which comprises the following steps of: carrying out manual marking and expansion processing on the starting and ending time of a sound fragment and a typical sound category; encoding the cry segments and the non-cry segments; performing random sampling on each coding unit; constructing a training data set; expanding a unit in the training set into multiple channels, performing feature extraction, and reconstructing the unit into a new multi-channel unit; carrying out model design; a complete pet record audio is input, and a two-dimensional structure is output; and generating a plurality of subsequences, and converting the subsequences into the starting and ending time of the cry fragment and the cry category to be output. Compared with the prior art, the method has the advantages that under the condition that a small amount of data is labeled, the sound is coded into the sequence, so that the model can adaptively recognize the sound segments with different lengths, and a sequence processing method is introduced into sound type recognition, so that the flexibility and accuracy of recognition are improved, and the interpretability of a result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of sound monitoring, especially a method for monitoring pet sounds. Background Art

[0002] With the rapid development of intelligent devices and health monitoring technologies, pet owners can use intelligent devices (such as pet collars, smart home systems, etc.) to monitor the sound activities of their pets in real time and understand their behaviors and health conditions. As a natural way for pets to express their physiological conditions, by analyzing the barks, breathing sounds, etc. of pets, key information about their health status can be provided, helping owners take better care of their pets and detect health problems in a timely manner. In particular, in aspects such as pet emotion recognition, health warning, and behavior tracking, the analysis of sound signals can provide more accurate feedback and control for pet owners.

[0003] Existing sound analysis technologies are mainly applied to human speech recognition or simple sound classification of certain animals. Usually, feature extraction methods such as Mel Frequency Cepstral Coefficients (MFCC) are used, and machine learning models (such as SVM, Convolutional Neural Network CNN, etc.) are combined for sound classification. In pet sound recognition, MFCC has been applied to a certain extent. However, due to the diversity of pet sounds and the large difference in their frequency ranges from human language, existing technologies often cannot fully capture health changes from pet sounds. Especially when dealing with health warning signals such as pet wheezing, barking, coughing, etc., the recognition accuracy has limitations.

[0004] In addition, existing technologies mainly focus on sound classification based on fixed time windows, and there are the following problems: First, the length of the input audio is fixed, resulting in the inability to fully express some information. Especially during long-term monitoring, important health or behavior information may be missed. Second, MFCC feature extraction is mainly designed for human speech and has poor adaptability to the frequency bands of pet barks, wheezing, coughing, etc., resulting in a decline in recognition performance in complex sound environments. In addition, existing data augmentation methods usually expand training data by upsampling, downsampling, or noise injection. However, these methods do not fully consider factors such as sound frequency changes, duration, and background noise, resulting in high homogeneity of the augmented data and difficulty in improving the generalization ability of the model. Moreover, when dealing with continuously changing sound signals, traditional methods often use truncation or zero-padding, resulting in damage to the effectiveness and continuity of the sound signals and affecting the accuracy of health monitoring.

[0005] For example, a Chinese invention patent with the patent publication number CN116230000A, applicant Suqian University, and title "A Blind Source Separation Method for Speech Noise Reduction" discloses the following content: The present invention discloses a blind source separation method for speech noise reduction in the technical field of blind source separation, including performing a first blind source separation operation on unknown strong interference signals in the source mixed signal to obtain the parameter characteristics of the unknown strong interference signals. The present invention includes acquiring multi-channel audio data of the environment where the target sound source is located; separating the multi-channel audio data based on a trained separation model to obtain single-channel audio data; and using the single-channel audio data as the audio data of the target sound source. The present invention solves the problem of speech overlap when multiple speakers speak in the same time period, and can accurately segment the speech and content of different speakers speaking in the same time period. Among them, when the multi-channel audio data is convolved with a two-dimensional convolution kernel, a two-dimensional feature is obtained. The behavior of this two-dimensional feature is the number of array elements of the microphone array. After the two-dimensional feature is encoded by the encoder, the three-dimensional matrix can represent the first audio feature.

[0006] The above patent has significant limitations in the application of pet acoustic scenarios: its hardware architecture relies on a multi-channel microphone array, and the complex deployment mode conflicts with the core requirements of pet multi-scenario switching and device portability and lightweight. Secondly, the training data of existing pre-trained models is mainly based on the spectral characteristics of human voices, while the frequency characteristics of pet vocalizations are different from those of human voices, resulting in insufficient cross-species generalization ability of the models. At the same time, the existing methods also over-focus on static physical quantities such as noise frequency, power, and amplitude at the technical level, lacking attention to the dynamic change characteristics of acoustic signals in the time-frequency domain, which may lead to limited signal noise reduction robustness in complex sound fields. Summary of the Invention

[0007] In order to solve the above problems existing in the prior art, the present application provides a pet sound monitoring method.

[0008] To achieve the above technical effects, the specific solution of the present application is as follows:

[0009] A pet sound monitoring method includes the following steps:

[0010] Step 1. In the complete audio recording of the pet, manually annotate the start and end times of the call segments and the typical call categories, and perform pre- and post-extension processing on the call segments;

[0011] Step 2. Encode the call segments and non-call segments to generate encoding units, and the encoding units include α units, β units, and γ units;

[0012] Step 3. Randomly sample each encoding unit to generate multiple random encoding units;

[0013] Step 4. Set the total number of training units to batch_num * N randomly encoded units, where batch_num represents the number of randomly encoded units in each training batch during model training, N is the total number of batches, and N is a positive integer. For the α units and β units in each training batch, extract the front-region α units, middle-region α units, and rear-region α units of the α units, and the front-region β units, middle-region β units, and rear-region β units of the β units respectively;

[0014] Step 5. Expand each randomly encoded unit in the training units into multiple channels, extract features from different channels of each randomly encoded unit, and reconstruct them into new multi-channel encoded units, thus completing the construction of the training dataset;

[0015] Step 6. Design a model based on the above training dataset;

[0016] Step 7. Input the complete pet record audio, generate multiple input units of the model after data preprocessing, train the model by constructing a convolutional neural network, set conditional loop iteration to find the optimal model parameters, then evaluate the generalization ability of the model using the validation set samples, and finally use the model with the highest validation accuracy as the optimal model for prediction, and output the two-dimensional structure D;

[0017] Step 8. Perform post-processing operations on the two-dimensional structure D output by the model to generate multiple subsequences, and each subsequence represents a predicted call segment;

[0018] Step 9. Convert the subsequences into the start and end times and call categories of the call segments for output.

[0019] Further, among the typical call categories marked in Step 1, the single-feature type calls refer to calls with a single acoustic feature, and the compound type calls are composed of combinations of multiple single-feature type calls. The single-feature type calls are divided into call type A and call type B, and the compound type calls are composed of segments of call type A and call type B, which are called call type C. In addition, all non-call segments are called non-call type Z.

[0020] Further, the front and rear expansions in Step 1 mean that for all call segments, the W sampling points before the starting point are called the front region, the W sampling points after the ending point are called the rear region, and the length of the call segment is the middle region, where W is the length of the unit during subsequent sampling.

[0021] Further, in Step 2, encode the call segments and non-call segments from a single call type. Encode call type A as α units, encode call type B as β units, and encode non-call segments as γ units.

[0022] Further, step 3 specifically includes:

[0023] The specific sampling rule for random sampling is as follows: for all non-call types Z, randomly extract the entire segment of each coding unit; for the segments of call type A and call type B of a single feature type call, for each call category, according to the quantity ratio of front_ratio:mid_ratio:back_ratio = x:y:z, randomly select the sampling starting point in the front, middle, and back regions of each coding unit respectively, and perform random sampling with length W.

[0024] Further, in step 4, for the random coding units in each batch during model training, the extraction ratios of α units, β units, and γ units are 1:1:1.

[0025] Further, step 4 specifically includes:

[0026] S4.1 Divide each batch into n sub-batches;

[0027] S4.2 Use random sampling without replacement within each sub-batch; use random sampling with replacement between sub-batches;

[0028] S4.3 Whenever entering the first sub-batch of each batch, initially, the number of units of each random coding unit is α unit = a, β unit = b, γ unit = c, satisfying and a < b < c, where batch_num represents the number of random coding units in each training batch during model training, and n is the number of sub-batches; when entering the second sub-batch, add 1 to the number of α units, add 1 to the number of β units, and subtract 2 from the number of γ units, and the number of each unit becomes α unit = a + 1, β unit = b + 1, γ unit = c - 2; follow this rule to cycle from the first sub-batch to the nth sub-batch in sequence. Thus, when each batch is passed through, the total number of α units is: The total number of β units is: The total number of γ units is c + (c - 2) + … + [c - 2*(n - 1)] = (c - n + 1)*n; finally, the extraction quantity ratio of α unit:β unit:γ unit is (c - n + 1)*n; where the range of the divided sub-batch n is: between 1 / 64 of the number of random coding units in each batch and 1 / 32 of the number of random coding units in each batch.

[0029] Further, in step 5, for each random coding unit, copy and expand the original single-channel random coding unit into multiple channels, construct different feature maps by combining different window functions, spectral estimations, and filters on each channel, and reconstruct it into a new multi-channel coding unit.

[0030] Further, in step 6, the constructed training dataset is divided according to the ratio of training set: validation set = 8:2 and then input into a pre-defined convolutional neural network;

[0031] (1) The number of network feature maps adopts a gradually decreasing design pattern.

[0032] (2) The initial feature extraction layer of the model is designed as "convolutional layer -> convolutional layer". In the first two layers, the "convolutional layer" is used and the convolutional kernel size is set to 3, and the stride is set to 1 to slide and extract the feature maps of the call segments. The model is trained through the constructed convolutional neural network, and the conditional loop iteration is set to find the optimal model parameters. Then, the generalization ability of the model is evaluated using the validation set samples, and finally the optimal model is output.

[0033] Further, in step 7 above, the data preprocessing method is to perform Fourier transform on the input audio data in the time domain in a sliding window manner to generate frequency domain data. For the frequency domain data of a single window, based on the vocal frequency range obtained by statistical analysis of the frequency range of the effective calls of the current pet, selection is performed, the data of multiple windows are accumulated, spliced to generate single-channel data, the single-channel data is copied and expanded into multiple channels, different feature extractions are performed, and k input units are generated for the entire audio segment. The optimal model is used for prediction, and a two-dimensional structure data D of k×d is output, where k represents the number of input units, and d is 3, representing the probabilities of the current input units being α, β, and γ respectively.

[0034] Further, in step 8 above, the post-processing operation specifically includes:

[0035] S8.1 Output the predicted input unit information for each row in the two-dimensional data D: Set a probability threshold, and compare the predicted probabilities of the α unit, β unit, and γ unit in each row with the probability threshold in turn. If the probability value of the α unit is greater than or equal to the probability threshold, the prediction result of this row is the α unit. If the probability value of the β unit is greater than or equal to the probability threshold, the prediction result of this row is the β unit. If the probability values of both the α and β units are less than the probability threshold, the prediction result of this row is the γ unit; After the two-dimensional data D goes through the above steps, the output unit sequence N will be obtained, and its size is k×1, including the α unit, β unit, and γ unit;

[0036] S8.2 Iteratively obtain the consecutive non-γ units in the unit sequence N, and record the consecutive non-γ units as the subsequence a. After traversing the unit sequence N, the subsequences a1 - a will be obtained n . For the subsequences a1 - a n perform a deletion operation;

[0037] Specifically, a deletion threshold is set. If the length of the subsequence is greater than the deletion threshold, the subsequence is retained; otherwise, the subsequence is deleted. After performing the deletion operation on the subsequence a1 - a n the subsequence b1 - b will be obtained n .

[0038] S8.3 performs a concatenation operation on the subsequence b1 - b n .

[0039] Specifically, a concatenation threshold is set, and the interval lengths between subsequences in b1 - b n are iteratively obtained. If the interval length is less than or equal to the concatenation threshold, the two subsequences are concatenated to obtain a longer subsequence. If the interval length is greater than the concatenation threshold, no modification is made. After the above steps, the subsequence c1 - c n will be obtained

[0040] Furthermore, in step 9 above, the specific method for outputting the start and end times of the call segment by converting the subsequence is as follows: The start time corresponding to the first input unit in the subsequence is used as the start time of the call segment, and the end time corresponding to the last input unit is used as the end time of the call segment. The method for outputting the call type by converting the subsequence is as follows: For the calls of types A, B, and C respectively, the proportions of α units and β units are statistically calculated. According to the statistical results of the proportions of each unit in the three types of calls, a first classification threshold cls_thr1, a second classification threshold cls_thr2, and a third classification threshold cls_thr3 are set to carry out multi-level call type judgment. The multi-level call type judgment method is as follows: For the subsequence, first judge whether the proportion of its α units is greater than cls_thr1. If it is greater, further judge between call types A and C. If it is less, further judge between call type B and non-call type Z. For the judgment between call types A and C, calculate the proportion of β units in the first m input units of the subsequence. If the proportion is greater than cls_thr2, output as call type C; otherwise, output as call type A. For the judgment between call type B and non-call type Z, calculate the proportion of β units in the subsequence. If the proportion is greater than cls_thr3, output as call type B

[0041] The above solution of the present invention has the following advantages compared with the existing technology

[0042] 1. The present invention provides a pet sound monitoring method. In daily life, the health status of pets is divided into three types: healthy, slightly uncomfortable and severely uncomfortable. The corresponding calls can also be divided into healthy calls, normal calls mixed with uncomfortable calls and uncomfortable calls. In this case, the method disclosed in the present invention can accurately classify different call states, thereby identifying different health states. Compared with the prior art, the present invention, with a small amount of data annotation, encodes the sound into a sequence, so that the model can adaptively identify sound fragments of different lengths, and introduces a sequence processing method in the sound type recognition, which not only improves the flexibility and accuracy of recognition, but also increases the interpretability of the results. At the same time, the present invention effectively optimizes model training through a data diversity balanced sampling strategy, so that it has higher recognition accuracy and stronger generalization ability with limited manpower and time costs. Overall, the method can quickly and accurately identify the vocal characteristics of pets in different health states, providing a convenient and efficient solution for pet health monitoring research.

[0043] 2. In the sound type encoding step disclosed in the present invention, the sound in the audio is mainly frame-coded. This will not change the type corresponding to the audio, and provides a basis for identifying sounds of different lengths. For a sound audio, it can be identified as an audio coding sequence, so that the recognizable audio length is more flexible. In addition, encoding the audio into a sequence can make the sequence more flexible, such as deleting and modifying, so as to increase a certain degree of interpretability while improving the subsequent recognition accuracy.

[0044] 3. The data diversity balanced sampling strategy disclosed in the present invention can effectively improve data diversity while controlling the impact of noise on subsequent model training. Compared with the problems of insufficient data utilization due to downsampling and relatively similar data increased by upsampling, partition sampling can obtain more diverse data while making full use of the data. Before inputting the model, compared with the purely random batch division of the data, by dividing a batch into multiple sub-batches, and using sampling with replacement between sub-batches and sampling without replacement within sub-batches, it is possible to maintain data diversity while controlling the proportion of data within the batch. This can effectively control the direction of model training and optimize the model training process, thereby reducing the oscillation phenomenon that may occur during model training and improving the generalization ability of the model.

[0045] 4. The hybrid sequence processing strategy disclosed in the present invention is a further processing based on the encoding strategy, which can achieve an effect similar to data denoising. By setting the probability threshold, ambiguous units can be converted into noise, thereby effectively improving the recognition accuracy. And the processing based on the deletion threshold can further filter out the noise interspersed in the sound segment, achieving an effect similar to denoising and ensuring the "validity" of the recognition segment. The subsequent connection threshold strategy can also make the recognized sound segment more complete on the basis of further filtering out the noise segment, thereby avoiding misclassification or omission of information and maintaining the "integrity" of the segment.

[0046] 5. The sound segment classification method disclosed in the present invention is mainly implemented through hierarchical classification. By counting the proportion of a certain type of sound category unit in the sound encoding sequence, the type of the sound segment can be initially judged. And by counting the proportion of another type of sound category unit, the supplement and correction of the initial judgment can be realized. Compared with the one-size-fits-all output of the segment type result by the model, hierarchical analysis can effectively decouple the sequence category information, providing interpretability for the final recognition result while improving the recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a schematic diagram of the overall method flow of this application.

[0048] Figure 2 It is a schematic diagram of the process of step 3.

[0049] Figure 3 It is a schematic diagram of the corresponding process from step 4 to step 5.

[0050] Figure 4 It is a schematic diagram of the model structure corresponding to step 6.

[0051] Figure 5 It is a schematic diagram of the corresponding process from step 7 to step 8.

[0052] Figure 6 It is a schematic diagram of the judgment process of the call category in step 9. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. Usually, the components of the embodiments of this application described and illustrated here can be arranged and designed in various different configurations.

[0054] Accordingly, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but merely represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the scope of protection of the present application.

[0055] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0056] In the description of the present application, it should be noted that the orientation or positional relationship indicated by terms such as "upper", "vertical", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the product of this application is usually placed during use, or the orientation or positional relationship commonly understood by those skilled in the art. It is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present application. In addition, terms such as "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0057] In the description of the present application, it should also be noted that unless otherwise clearly specified and defined, terms such as "set", "installed", "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.

[0058] Embodiment 1

[0059] A pet sound monitoring method includes the following steps:

[0060] Step 1. In the complete audio recording of the pet, manually mark the start and end times of the call segments and the typical call categories, and perform front and back extension processing on the call segments;

[0061] Step 2. Encode the call segments and non-call segments to generate encoding units, where the encoding units include α units, β units, and γ units;

[0062] Step 3. Randomly sample the different encoding units defined above to generate multiple random encoding units;

[0063] Step 4. When constructing the training data set, the total number of training units is set to batch_num*N random coding units, where batch_num represents the number of random coding units in each training batch during model training (e.g., batch_num=640), N is the total number of control training units, and N is a positive integer. For the α units and β units in each training batch of the subsequent model, the front area α units, middle area α units, and rear area α units of the α units, and the front area β units, middle area β units, and rear area β units of the β units are extracted respectively;

[0064] Step 5. Expand each random coding unit in the training unit into multiple channels, extract features from different channels of each random coding unit, and reconstruct them into new multi-channel coding units, thereby completing the construction of the training data set;

[0065] Step 6. Design a model based on the above training data set;

[0066] Step 7. Input the complete pet recording audio, generate multiple input units of the model after data preprocessing, train the model by building a convolutional neural network, set conditional loop iterations to find the optimal model parameters, and then use the validation set samples to evaluate the generalization ability of the model. Finally, use the model with the highest validation accuracy as the optimal model for prediction, and output the two-dimensional structure D.

[0067] Step 8. Perform post-processing operations on the two-dimensional structure D output by the model to generate multiple subsequences, each subsequence representing a predicted call segment;

[0068] Step 9: Convert the subsequence into the start and end time of the call segment and the call category for output.

[0069] Compared with the prior art, the present invention, with a small amount of data annotation, encodes the sound into a sequence, so that the model can adaptively identify sound segments of different lengths, and introduces a sequence processing method in the sound type recognition, which not only improves the flexibility and accuracy of recognition, but also increases the interpretability of the results. At the same time, the present invention effectively optimizes model training through a data diversity balanced sampling strategy, so that it has higher recognition accuracy and stronger generalization ability with limited manpower and time costs. Overall, this method can quickly and accurately identify the vocal characteristics of pets in different health states, providing a convenient and efficient solution for pet health monitoring research.

[0070] Example 2

[0071] A pet sound monitoring method comprises the following steps:

[0072] Step 1. In the complete audio recording of the pet, manually annotate the start and end times of the barking segments and the typical barking categories, and perform pre- and post-expansion processing on the barking segments;

[0073] Step 2. Encode the barking segments and non-barking segments to generate encoding units, which include α units, β units, and γ units;

[0074] Step 3. Randomly sample the different encoding units defined above to generate multiple random encoding units;

[0075] Step 4. When constructing the training dataset, set the total number of training units to batch_num * N random encoding units, where batch_num represents the number of random encoding units in each training batch during model training (for example, batch_num = 640), and N is used to control the total number of training units. N is a positive integer. For the α units and β units in each training batch for training the subsequent model, extract the front-region α units, middle-region α units, and rear-region α units of the α units, and the front-region β units, middle-region β units, and rear-region β units of the β units respectively;

[0076] Step 5. Expand each random encoding unit in the training units into multiple channels, extract features from different channels of each random encoding unit, and reconstruct them into new multi-channel encoding units, thus completing the construction of the training dataset;

[0077] Step 6. Design a model based on the above training dataset;

[0078] Step 7. Input the complete audio recording of the pet, generate multiple input units of the model after data preprocessing, train the model by constructing a convolutional neural network, set conditional loop iteration to find the optimal model parameters, then evaluate the generalization ability of the model using the validation set samples, and finally use the model with the highest validation accuracy as the optimal model for prediction, and output the two-dimensional structure D;

[0079] Step 8. Perform post-processing operations on the two-dimensional structure D output by the model to generate multiple subsequences, and each subsequence represents a predicted barking segment;

[0080] Step 9. Convert the subsequences into the start and end times and barking categories of the barking segments for output.

[0081] The typical call categories marked in Step 1 are divided into single - feature - type calls and composite - type calls. Among them, single - feature - type calls refer to calls with single, clear, and stable acoustic features, and the acoustic features include frequency, pitch, and amplitude. Composite - type calls are composed of multiple single - feature - type calls combined according to certain rules, and these features can also change and intertwine to form more complex acoustic patterns. Based on the differences in acoustic features, single - feature - type calls are divided into Call Type A and Call Type B. Composite - type calls are composed of segments of Call Type A and Call Type B and are called Call Type C. In addition, all non - call segments are called Non - call Type Z.

[0082] For example, Call Type A is a normal call, including ordinary calls and functional calls in normal physiological states. Functional calls in normal physiological states include alert calls, courtship calls, etc. Call Type B is a discomfort call, including calls due to injuries, illness, hunger, fatigue, anxiety, or other abnormal physiological states. Call Type C is a normal call mixed with discomfort calls.

[0083] The front - and - back extension of Step 1 means that for all call segments, the W sampling points before the starting point are called the front region, the W sampling points after the ending point are called the back region, and the length of the call segment is the middle region, where W is the length of the unit in subsequent sampling.

[0084] In Step 2, the call segments and non - call segments from a single call type are encoded. Specifically: Call Type A is encoded as α units, Call Type B is encoded as β units, and non - call segments are encoded as γ units.

[0085] Step 3 specifically includes:

[0086] The specific sampling rules for random sampling are as follows: for all non-call types Z, the entire coding unit of each segment is randomly selected; for the segments of call type A and call type B of a single feature type call, for each call category, according to the quantity ratio of front_ratio:mid_ratio:back_ratio = x:y:z, the sampling starting points are randomly selected in the front, middle, and back regions of each coding unit respectively, and random sampling is carried out with a length of W (i.e., for call category A: α units in the front region / α units in the middle region / α units in the back region; for call category B: β units in the front region / β units in the middle region / β units in the back region). When randomly selecting sampling starting points in the front, middle, and back regions according to the quantity ratio of x:y:z, the proportions of x and z are 10% to 20% respectively, and the proportion of y is 60% to 80%. Optionally, when randomly selecting sampling starting points in the front, middle, and back regions according to the quantity ratio of x:y:z, the quantity ratio of x:y:z is 1:8:1. The advantages are: (1) The extraction of call segment edge information (i.e., the front region / the back region: including the conversion information from background / noise to call segment) helps to increase the diversity of data, thereby improving the generalization of the subsequent generation model; (2) Setting a smaller extraction ratio for the front / back regions (i.e., setting smaller extraction values for x and z) is to reduce the impact of silence and noise on the subsequent model training.

[0087] In step 4, for the random coding units in each batch during model training, the extraction ratios of α units (including α units in the front region / α units in the middle region / α units in the back region), β units (including β units in the front region / β units in the middle region / β units in the back region), and γ units are 1:1:1.

[0088] Step 4 specifically includes:

[0089] S4.1 Divide each batch (e.g., batch_num = 640) into n sub-batches

[0090] S4.2 In each sub-batch, random sampling without replacement is used to ensure that each type of unit in the sub-batch is not repeated; random sampling with replacement is used between sub-batches to increase the diversity of unit combinations between different sub-batches;

[0091] S4.3 Whenever entering the first sub-batch of each batch, initially, the number of units of each random coding unit is α unit = a, β unit = b, γ unit = c, satisfying And a < b < c, where batch_num represents the number of random coding units in each training batch during model training, and n is the number of sub-batches; when entering the second sub-batch, the number of α units is incremented by 1, the number of β units is incremented by 1, and the number of γ units is decremented by 2. The quantity of each type of unit becomes α unit = a + 1, β unit = b + 1, and γ unit = c - 2; following this pattern, it cycles from the first sub-batch to the nth sub-batch in sequence. Thus, when each batch (i.e., n sub-batches) has passed, the total quantity of α units is: The total quantity of β units is: The total quantity of γ units is c+(c - 2)+…+[c - 2*(n - 1)]=(c - n + 1)*n; finally, the extraction quantity ratio of α unit:β unit:γ unit is (c - n + 1)*n; by presetting the split sub-batch n, initializing the number of α units a, initializing the number of β units b, and initializing the number of γ units c, the quantity ratio of α units, β units, and γ units entering each batch is approximately 1:1:1. Among them, the range of the split sub-batch n is: between 1 / 64 of the number of random coding units in each batch and 1 / 32 of the number of random coding units in each batch. E.g., set n = 10, a = 16, b = 17, c = 31, and the finally obtained extraction quantity ratio of α unit:β unit:γ unit is 205:215:220≈1:1:1, and the current batch_num = 640.

[0092] In step 5, for each random coding unit, in order to extract its features from different dimensions, the original single-channel random coding unit is copied and expanded into multiple channels. Different feature maps are constructed on each channel by combining, including but not limited to, different window functions, spectral estimation, and filters, and it is reconstructed into a new multi-channel coding unit. For example, each unit can be reconstructed into a three-channel unit. On the first channel of each unit, the Hanning window and Welch method are used to estimate the power spectrum and take the logarithm, a specific frequency band is selected according to the characteristics of the research object, and it is filtered with a Gaussian filter and normalized; on the second channel of each unit, the Hamming window and Welch method are used to estimate the power spectrum and take the logarithm, a specific frequency band is selected according to the characteristics of the research object, and it is filtered with a Gaussian filter and normalized; on the third channel of each unit, the Hann window and Welch method are used to estimate the power spectrum and take the logarithm, a specific frequency band is selected according to the characteristics of the research object, and it is filtered with a Savitzky Golay filter and normalized.

[0093] In step 6, the constructed training dataset (i.e., the data size is [M, 96, 128, C], where M is the number of samples and C is the number of channels) is divided according to the ratio of training set:validation set = 8:2 and then input into a pre-defined convolutional neural network (CNN).

[0094] Specifically, (1) the number of network feature maps adopts a gradually decreasing design pattern. The traditional neural network has an increasing pattern of the number of feature maps, resulting in too many parameters in the fully connected layer, which causes the parameter update in this network structure to be mainly concentrated in the fully connected layer (i.e., the feature fusion part). Therefore, the proportion of the number of parameter updates in the convolutional layer of the feature extraction part is relatively small, which is not conducive to the model's extraction of call features. (2) The combination of "convolutional layer -> max pooling layer" in the initial feature extraction layer of the model is changed to the design of "convolutional layer -> convolutional layer". Since the size of the call segment feature map input to the network is 96*128 (i.e., the feature map size is relatively small), if the "convolutional layer -> max pooling layer" construction method is directly used in the first two layers of the network, some important features will be severely compressed after passing through the "max pooling layer", resulting in the loss of call feature information. Therefore, "convolutional layers" are used in both the first two layers, and the convolutional kernel size is set to 3 and the stride is set to 1 to slide and extract the feature map of the call segment, so as to obtain more detailed information in the initial stage. The model is trained through the constructed convolutional neural network, and the conditional loop iteration is set to find the optimal model parameters, and then the generalization ability of the model is evaluated using the validation set samples, and finally the optimal model is output.

[0095] In step 7, the data preprocessing method is as follows: the input audio data is Fourier-transformed in the time domain in a sliding window manner to generate frequency domain data. For the frequency domain data of a single window, selection is made based on the vocal frequency range obtained through statistical analysis of the frequency range of the current pet's effective calls. The data of multiple windows are accumulated and spliced to generate single-channel data. The single-channel data is copied and extended to multiple channels for different feature extractions. For example: Gaussian filter filtering and normalization are performed on the first channel, logarithmic operation is performed on the second channel, and Savitzky-Golay filter filtering and normalization are performed on the third channel. The multiple channels with different processing methods are integrated to obtain the input unit of the model. The entire audio will generate k input units through the above method. The optimal model is used for prediction, and a two-dimensional structure data D of k×d will be output, where k represents the number of input units and d is 3, representing the probabilities that the current input unit is α, β, and γ respectively.

[0096] In step 8, the post-processing operation specifically includes:

[0097] S8.1 Output the predicted input unit information for each row in the two-dimensional data D. Considering that more accurate prediction results can be obtained for the call α unit and β unit, the design method is as follows: Set a probability threshold, and compare the predicted probabilities of the α unit, β unit, and γ unit in each row with the probability threshold in turn. If the probability value of the α unit is greater than or equal to the probability threshold, the prediction result of this row is the α unit; if the probability value of the β unit is greater than or equal to the probability threshold, the prediction result of this row is the β unit; if the probability values of both the α and β units are less than the probability threshold, the prediction result of this row is the γ unit. After the two-dimensional data D goes through the above steps, the output unit sequence N will be obtained, which has a size of k×1 and contains the α unit, β unit, and γ unit.

[0098] S8.2 Iteratively obtain the consecutive non-γ units in the unit sequence N, and denote the consecutive non-γ units as the subsequence a. After traversing the unit sequence N, the subsequences a1 - a will be obtained. n . To obtain more accurate call segment information, perform deletion operations on the subsequences a1 - a. n Perform deletion operations.

[0099] Set a deletion threshold. If the length of the subsequence is greater than the deletion threshold, retain the subsequence; otherwise, delete the subsequence. After performing deletion operations on the subsequences a1 - a, the subsequences b1 - b will be obtained. n n .

[0100] S8.3 To smooth the noise segments mispredicted by the model, perform concatenation operations on the subsequences b1 - b. n Perform concatenation operations.

[0101] Specifically, set a concatenation threshold, and iteratively obtain the interval length between subsequences in b1 - b. n If the interval length is less than or equal to the concatenation threshold, concatenate these two subsequences to obtain a longer subsequence; if the interval length is greater than the concatenation threshold, no modification is made. For example, if the interval length between subsequence b1 and b2 is less than the concatenation threshold, use the starting position of subsequence b1 as the starting position of the new subsequence c, and use the ending position of subsequence b2 as the ending position of the new subsequence c. This step concatenates subsequences b1 and b2 into c, and then continues to determine whether the interval length between c and b3 meets the concatenation threshold. After the above steps, the subsequences c1 - c will be obtained. n .

[0102] Further, in the above step 9, the specific method for outputting the start and end times of converting the subsequence into a call segment is as follows: the start time corresponding to the first input unit in the subsequence is used as the start time of the call segment, and the end time corresponding to the last input unit is used as the end time of the call segment; the method for outputting the call category by converting the subsequence is as follows: for the calls of three types A, B, and C, the proportions of α units and β units are respectively statistically analyzed. According to the statistical results of the proportions of each unit in the three types of calls, the first classification threshold cls_thr1, the second classification threshold cls_thr2, and the third classification threshold cls_thr3 are set to carry out multi-level call type judgment. The multi-level call type judgment method is as follows: for the subsequence, first judge whether the proportion of its α units is greater than cls_thr1. If it is greater, further judge the call types A and C. If it is less, further judge the call types B and the non-call type Z. For the call types A and C, calculate the proportion of β units in the first m input units of the subsequence. If the proportion is greater than cls_thr2, output that it is the call type C, otherwise output that it is the call type A; for the call types B and the non-call type Z, calculate the proportion of β units in the subsequence. If the proportion is greater than cls_thr3, output that it is the call type B.

Claims

1. A pet sound monitoring method, characterized in that, It includes the following steps: Step 1. In the complete audio recording of the pet, manually annotate the start and end times and typical call categories of the call segments, and perform pre- and post-extension processing on the call segments; Step 2. Encode the call segments and non-call segments to generate encoding units, where the encoding units include α units, β units, and γ units; Step 3. Randomly sample each encoding unit to generate multiple random encoding units; Step 4. Set the total number of training units to batch_num * N random encoding units, where batch_num represents the number of random encoding units in each training batch during model training, N is the total number of batches, and N is a positive integer. For the α units and β units in each training batch, extract the front-region α units, middle-region α units, and rear-region α units of the α units, and the front-region β units, middle-region β units, and rear-region β units of the β units respectively; Step 5. Expand each random encoding unit in the training units into multiple channels, extract features from different channels of each random encoding unit, and reconstruct them into new multi-channel encoding units, thereby completing the construction of the training dataset; Step 6. Design a model based on the above training dataset; Step 7. Input the complete audio recording of the pet, generate multiple input units of the model after data preprocessing, train the model by constructing a convolutional neural network, set conditional loop iteration to find the optimal model parameters, then evaluate the generalization ability of the model using the validation set samples, and finally use the model with the highest validation accuracy as the optimal model for prediction, and output the two-dimensional structure D; Step 8. Perform post-processing operations on the two-dimensional structure D output by the model to generate multiple subsequences, and each subsequence represents a predicted call segment; Step 9. Convert the subsequences into the start and end times and call categories of the call segments for output.

2. The pet sound monitoring method according to claim 1, characterized in that Among the typical call categories annotated in Step 1, there are single-feature type calls and composite type calls. A single-feature type call refers to a call with a single acoustic feature, and a composite type call is composed of multiple single-feature type calls. The single-feature type calls are divided into call type A and call type B, and the composite type call is composed of segments of call type A and call type B, which is called call type C. In addition, all non-call segments are called non-call type Z; The pre- and post-extension in Step 1 means that for all call segments, the W sampling points before the starting point are called the front region, the W sampling points after the ending point are called the rear region, and the length of the call segment is the middle region, where W is the length of the unit during subsequent sampling.

3. The method for monitoring pet sounds according to claim 2, characterized in that, In Step 2, encode the call segments and non-call segments from a single call type. Encode call type A as an α unit, encode call type B as a β unit, and encode non-call segments as γ units.

4. The pet sound monitoring method according to claim 3, wherein The specific content of Step 3 includes: The specific sampling rules for random sampling are as follows: for all non-call types Z, randomly extract each entire coding unit; for the segments of call type A and call type B of a single feature type call, for each call category, randomly select sampling starting points in the front, middle, and back regions of each coding unit according to the quantity ratio of front_ratio:mid_ratio:back_ratio = x:y:z, and perform random sampling with length W respectively.

5. A pet sound monitoring method according to claim 4, characterized in that, In step 4, for the random coding units in each batch during model training, the extraction ratios of α units, β units, and γ units are 1:1:1; the specific steps of step 4 include: S4.1 Divide each batch into n sub-batches; S4.2 Use random sampling without replacement within each sub-batch; use random sampling with replacement between sub-batches; S4.3 Whenever entering the first sub-batch of each batch, initially, the number of each type of randomly encoded unit is α unit = a, β unit = b, γ unit = c, satisfying and a < b < c, where batch_num represents the number of randomly encoded units in each training batch during model training, and n is the number of sub-batches; when entering the second sub-batch, the number of α units is increased by 1, the number of β units is increased by 1, and the number of γ units is decreased by 2. The number of each type of unit becomes α unit = a + 1, β unit = b + 1, γ unit = c - 2; following this rule, it cycles from the first sub-batch to the nth sub-batch in sequence. Thus, when each batch is passed, the total number of α units is: The total number of β units is: The total number of γ units is c + (c - 2) + … + [c - 2*(n - 1)] = (c - n + 1)*n; finally, the extraction quantity ratio of α unit : β unit : γ unit is (c - n + 1)*n; where the range of the divided sub-batch n is: between 1 / 64 and 1 / 32 of the number of randomly encoded units in each batch.

6. A method for monitoring pet sounds according to claim 5, characterized in that, In step 5, for each random coding unit, copy and expand the original single-channel random coding unit into multiple channels, construct a feature map by combining a window function, spectral estimation, and a filter on each channel, and reconstruct it into a new multi-channel coding unit.

7. A pet sound monitoring method according to claim 6, characterized in that, In step 6, divide the already constructed training dataset into a training set and a validation set according to a ratio of 8:2, and input it into a pre-defined convolutional neural network; The number of network feature maps adopts a design pattern of gradually decreasing; The initial feature extraction layer of the model is designed as "convolution layer -> convolution layer". In the first two layers, both use the "convolution layer" and set the convolution kernel size to 3 and the stride to 1 to slide and extract the feature map of the call segment. Train the model through the constructed convolutional neural network, set the conditional loop iteration to find the optimal model parameters, then evaluate the generalization ability of the model with the validation set samples, and finally output the optimal model.

8. A pet sound monitoring method according to claim 7, characterized in that, In step 7 above, the data preprocessing method is to perform Fourier transform on the input audio data in the time domain in a sliding window manner to generate frequency domain data. For the frequency domain data of a single window, select it based on the vocal frequency range obtained by statistical analysis of the frequency range of the effective calls of the current pet. Accumulate the data of multiple windows, splice them to generate single-channel data, copy and expand the single-channel data into multiple channels, perform different feature extractions, generate k input units for the entire audio segment, use the optimal model for prediction, and output the two-dimensional structure data D of k×d, where k represents the number of input units and d is 3, respectively representing the probabilities that the current input unit is α, β, and γ.

9. The pet sound monitoring method according to claim 8, characterized in that, In step 8 above, the post-processing operations specifically include: S8.1 Output the input unit information predicted for each row in the two-dimensional data D: Set a probability threshold, and compare the predicted probabilities of the α unit, β unit, and γ unit in each row with the probability threshold in turn. If the probability value of the α unit is greater than or equal to the probability threshold, the prediction result of this row is the α unit. If the probability value of the β unit is greater than or equal to the probability threshold, the prediction result of this row is the β unit. If the probability values of both the α and β units are less than the probability threshold, the prediction result of this row is the γ unit; after the two-dimensional data D goes through the above steps, the output unit sequence N will be output, with a size of k×1, containing the α unit, β unit, and γ unit; S8.2 Iteratively obtain consecutive non-γ cells in the cell sequence N, denote the consecutive non-γ cells as subsequence a, and subsequences a1 - a will be obtained after traversing the cell sequence N n , and perform a deletion operation on subsequences a1 - a n ; Set a deletion threshold. If the length of the subsequence is greater than the deletion threshold, retain the subsequence; otherwise, delete the subsequence. After performing the deletion operation on the subsequence a1 - a n the resulting subsequence will be b1 - b n ; S8.3 Connect the subsequences b1 - b n Perform a connection operation; Set a connection threshold and iteratively obtain b1 - b n The interval length between subsequences; if the interval length is less than or equal to the connection threshold, then connect these two subsequences to obtain a longer subsequence. If the interval length is greater than the connection threshold, no modification is made. After the above steps, subsequences c1 - c will be obtained n .

10. A pet sound monitoring method according to claim 9, characterized in that, In the above step 9, the specific method for outputting the start and end times of converting the subsequence into a call segment is: Use the start time corresponding to the first input unit in the subsequence as the start time of this call segment, and the end time corresponding to the last input unit as the end time of this call segment; the method for outputting the call category by converting the subsequence is: For the three types of calls A, B, and C, respectively, count the proportions of the α unit and β unit. According to the statistical results of the proportions of each unit in the three types of calls, set the first classification threshold cls_thr1, the second classification threshold cls_thr2, and the third classification threshold cls_thr3 to carry out multi-level call type judgment. The multi-level call type judgment method is: For the subsequence, first judge whether the proportion of its α unit is greater than cls_thr1. If it is greater, further judge the call types A and C. If it is less, further judge the call types B and the non-call type Z. For the call types A and C, calculate the proportion of the β unit in the first m input units of the subsequence. If the proportion is greater than cls_thr2, output that it is the call type C, otherwise output that it is the call type A; for the call types B and the non-call type Z, calculate the proportion of the β unit in the subsequence. If the proportion is greater than cls_thr3, output that it is the call type B.

Citation Information

Patent Citations

  • Blind source separation method applied to voice noise reduction

    CN116230000A