Crying sound classification prediction method and device, electronic equipment and storage medium

By acquiring and fusing infant audio sequence and historical behavior information, the problem of accuracy and inefficiency in the prediction of cry classification in the prior art is solved, and more efficient and accurate prediction results are achieved.

CN120015016APending Publication Date: 2025-05-16CHINA MOBILEHANGZHOUINFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311519522.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-14
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art has low accuracy and efficiency in predicting infant cry sound classification, single-modal methods are difficult to provide multi-angle prediction, and multi-modal methods have high computational complexity.

Method used

By obtaining the audio sequence and historical behavior information of the target user, cry type prediction and adjustment vector determination are performed separately, and the prediction results are fused to improve prediction accuracy and reduce calculation complexity.

Benefits of technology

The accuracy and efficiency of crying classification prediction are improved, and the calculation difficulty is reduced through the lightweight multimodal structure to achieve more efficient prediction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015016A_ABST
    Figure CN120015016A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and provides a cry classification prediction method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining an audio sequence and historical behavior information of a target user; performing cry type prediction based on the audio sequence to obtain a cry type prediction vector; determining an adjustment vector of the target user based on the historical behavior information; and fusing the cry type prediction vector and the adjustment vector to obtain a target prediction result. According to the cry classification prediction method provided by the embodiment of the invention, the audio sequence and the historical behavior information are respectively used as independent modal data, the cry type prediction vector is obtained through the cry type prediction of the audio sequence, and the adjustment vector is obtained through the adjustment mode of the historical behavior; therefore, the target prediction result is predicted by fusing the output of each mode, the adjustment of the cry type prediction vector is realized, the accuracy of cry classification prediction can be improved, and the efficiency of cry classification prediction can be improved through a lightweight multi-mode structure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a crying sound classification prediction method, device, electronic equipment and storage medium. Background Art

[0002] Crying is the most important means of information transmission for infants, which can reflect the current physiological and psychological needs of infants. For people who take care of infants, if they can understand the reason why infants cry, they can take effective countermeasures to meet the needs of infants. At present, the mainstream processing method of infant crying data in the industry is usually based only on audio data, using audio preprocessing features and deep neural networks to detect or classify crying. This method can achieve good performance in the task of infant crying detection, but the performance in the task of infant crying classification is not satisfactory.

[0003] Under the existing technology, the unimodal method using a single data source is relatively one-sided and difficult to provide predictions from multiple angles, resulting in low accuracy of cry classification predictions; moreover, the multimodal method may require the combined use of multiple neural networks to implement, and its computational difficulty will increase accordingly, resulting in low efficiency of cry classification predictions. Summary of the invention

[0004] The embodiments of the present invention provide a crying sound classification prediction method, device, electronic device and storage medium to solve the problems of low accuracy and efficiency of crying sound classification prediction.

[0005] In a first aspect, an embodiment of the present invention provides a crying sound classification prediction method, comprising:

[0006] Obtain the target user's audio sequence and historical behavior information;

[0007] Predicting the crying type based on the audio sequence to obtain a crying type prediction vector;

[0008] Based on the historical behavior information, determining an adjustment vector of the target user; the adjustment vector is used to supplement the crying type prediction vector with historical behavior;

[0009] The cry type prediction vector and the adjustment vector are fused to obtain a target prediction result.

[0010] In one embodiment, determining the adjustment vector of the target user based on the historical behavior information includes:

[0011] Determine the historical behavior interval time based on the historical behavior time of the historical behavior information;

[0012] Determining a probability parameter of a cause of the historical behavior based on the historical behavior interval;

[0013] Based on the probability parameter, the adjustment vector is determined.

[0014] In one embodiment, the step of fusing the cry type prediction vector and the adjustment vector to obtain a target prediction result includes:

[0015] Performing matrix transposition processing on the adjustment vector to obtain an adjustment transposed vector;

[0016] Performing a dot product calculation on the crying type prediction vector and the adjusted transposed vector to obtain a target prediction vector;

[0017] Based on the target prediction vector, determining a target probability value for each crying type in the target prediction vector;

[0018] A target prediction result is determined based on the target probability value of each cry type in the target prediction vector.

[0019] In one embodiment, determining the target prediction result based on the target probability value of each cry type in the target prediction vector includes:

[0020] Determining a maximum target probability value based on the target probability values ​​of each cry type in the target prediction vector;

[0021] Based on the maximum target probability value, a target prediction result is determined.

[0022] In one embodiment, the step of predicting the crying type based on the audio sequence to obtain a crying type prediction vector includes:

[0023] Performing audio preprocessing on the audio sequence to obtain at least one audio segment feature;

[0024] Inputting each of the audio segment features into a crying type prediction model to obtain at least one crying type vector of the segment output by the crying type prediction model; the crying type prediction model is obtained by training an audio neural network;

[0025] Based on the cry type vectors of each segment, a cry type prediction vector is determined.

[0026] In one embodiment, the performing audio preprocessing on the audio sequence to obtain at least one audio segment feature includes:

[0027] Segmenting the audio sequence to obtain at least one audio segment;

[0028] Perform spectrum conversion on each of the audio segments to obtain at least one audio segment feature.

[0029] In one embodiment, determining the crying type prediction vector based on the crying type vectors of each segment includes:

[0030] The cry type vectors of each segment are averaged to obtain the cry type prediction vector.

[0031] In a second aspect, an embodiment of the present invention provides a crying sound classification prediction device, comprising:

[0032] The acquisition module is used to obtain the audio sequence and historical behavior information of the target user;

[0033] A prediction module, used for predicting the crying type based on the audio sequence to obtain a crying type prediction vector;

[0034] An adjustment module, used for determining an adjustment vector of the target user based on the historical behavior information; the adjustment vector is used for supplementing the crying type prediction vector with historical behavior;

[0035] A fusion module is used to fuse the cry type prediction vector and the adjustment vector to obtain a target prediction result.

[0036] In a third aspect, an embodiment of the present invention provides an electronic device, including a processor and a memory storing a computer program, wherein when the processor executes the program, the crying sound classification prediction method described in the first aspect is implemented.

[0037] In a fourth aspect, an embodiment of the present invention provides a storage medium, which is a computer-readable storage medium and includes a computer program. When the computer program is executed by a processor, the crying classification prediction method described in the first aspect is implemented.

[0038] The crying classification prediction method, device, electronic device and storage medium provided in the embodiments of the present invention take the audio sequence and historical behavior information as independent modal data respectively, obtain the crying type prediction vector by predicting the crying type of the audio sequence, obtain the adjustment vector by the adjustment mode of the historical behavior, and then predict the target prediction result by fusing the outputs of each modality to achieve the adjustment of the crying type prediction vector, which can improve the accuracy of crying classification prediction, and can also reduce the calculation difficulty and improve the efficiency of crying classification prediction through a lightweight multi-modal structure. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0040] Figure 1 is a flow chart of a crying sound classification prediction method provided by an embodiment of the present invention;

[0041] Figure 2 is a schematic diagram of the framework of a crying sound classification prediction device provided by an embodiment of the present invention;

[0042] Figure 3 is a flow chart of the overall solution of the crying sound classification prediction method provided by an embodiment of the present invention;

[0043] Figure 4 It is a schematic diagram of functional modules of an embodiment of a crying sound classification prediction device of the present invention;

[0044] Figure 5 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0045] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution in the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. In the description of this specification, the description of the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the embodiment of the present invention. In this specification, the schematic representation of the above terms does not necessarily target the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, in the absence of mutual contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work belong to the scope of protection of the present invention.

[0046] The crying sound classification prediction method, device, electronic device and storage medium provided by the present invention are described in detail below in conjunction with the embodiments.

[0047] Figure 1 is a flow chart of a crying sound classification prediction method provided by an embodiment of the present invention;

[0048] Figure 2 Schematic diagram of the framework of the crying sound classification prediction device provided by an embodiment of the present invention.

[0049] Reference Figure 1, an embodiment of the present invention provides a crying sound classification prediction method, which may include:

[0050] Step 100, obtaining the audio sequence and historical behavior information of the target user;

[0051] Step 200, predicting the crying type based on the audio sequence to obtain a crying type prediction vector;

[0052] Step 300, determining an adjustment vector of the target user based on the historical behavior information;

[0053] Step 400: Fusing the cry type prediction vector and the adjustment vector to obtain a target prediction result.

[0054] It should be noted that infants usually refer to human infants under 12 months old. There are two main types of attachment behaviors: active behaviors (such as following or crawling) and signal behaviors (such as crying, laughing and babbling). Before newborns have the ability to take the initiative, they can only rely on signal behaviors to attract adult attention and convey information. Among signal behaviors, crying is a behavior that reflects the survival status of infants, and is usually related to negative states (such as hunger, sleepiness, pain, etc.). Therefore, the crying of infants is the most important signal of all signal behaviors. Based on the physiological characteristics of infants, their crying may be related to their experiences or behaviors in a certain period of time before. For example, infants are unlikely to cry because of hunger for a period of time after being fed, or they will not immediately become sleepy if they have enough sleep. If the baby has cried because of hunger and has not been fed, the probability that the next crying is due to hunger will increase.

[0055] It should be further explained that the crying sound classification prediction method provided in the embodiment of the present invention is implemented based on the crying sound classification prediction device, and the crying sound classification prediction method is mainly used for the classification prediction of infant crying to determine the cause of the infant crying and take corresponding measures according to the cause of the infant crying. Therefore, the embodiment of the present invention takes the crying sound classification prediction device as an example to describe the crying sound classification prediction method.

[0056] Specifically, refer to Figure 2It can be seen that the crying classification prediction device includes a data acquisition module, which is used to collect audio data of the target user. The data acquisition module can be a microphone or other audio sensor; the crying classification prediction device includes a behavior recording module, which is used to record the behavior information of the target user. The behavior recording module can record the behavior information of the target user by manually recording the behavior information through the user terminal, or by recording the historical crying type prediction results, and does not require the use of additional collection equipment. Therefore, the crying classification prediction device obtains the audio data of the target user through the data acquisition module, and obtains the historical behavior information of the target user through the behavior recording module. Furthermore, the crying classification prediction device converts the obtained audio data of the target user into a sequence to obtain the audio sequence of the target user.

[0057] Further, refer to Figure 2 It can be seen that the crying sound classification prediction device includes a classification prediction module, which is used to predict the crying sound type based on the audio sequence of the target user. Therefore, the crying sound classification prediction device predicts the crying sound type based on the audio sequence to obtain a crying sound type prediction vector.

[0058] Therefore, it can be understood that the crying classification prediction device performs audio preprocessing on the audio sequence of the target user to obtain at least one audio segment feature, and further inputs each audio segment feature into the crying type prediction model to obtain at least one segment crying type vector output by the crying type prediction model, and determines the crying type prediction vector based on each segment crying type vector.

[0059] It should be noted that the classification prediction module is a module for calculating the target user's cry type prediction vector. This module can be any algorithm for calculating the cry type prediction vector and is highly interchangeable. Specifically, the algorithm used by this module can be a single-modal or multi-modal method, and can use traditional algorithms, machine learning or deep learning models. The input of this module comes from the data acquisition module, which is all the data required by the algorithm used by this module. Therefore, the formula of the classification prediction module is described as:

[0060] f(X)=Y

[0061] Among them, X is the input data collected by the behavior recording module, Y is the crying type prediction vector, and f(·) is the specific algorithm used by the classification prediction module.

[0062] Further, refer to Figure 2It can be seen that the crying classification prediction device includes a historical behavior adjustment module, which is used to supplement the crying type prediction vector output by the classification prediction module with historical behavior. It should be noted that there are two main categories of historical information related to infants: crying type and behavior related to infants. Among them, the crying type can be adjusted according to the purpose, for example, to distinguish the needs of infants (such as hunger, sleepiness, skin discomfort, etc.), or to judge the pain level of infants; and the behavior records related to infants also need to change with the purpose. For example, when used to distinguish the needs of infants, it is necessary to record the feeding, sleeping, bathing and diaper changing of infants, etc., and when used to judge the pain level, it is necessary to consider the frequency of crying and the pain level over a period of time.

[0063] Therefore, the crying classification prediction device determines the adjustment vector of the target user based on the historical behavior information of the target user, wherein the adjustment vector of the target user is used to supplement the crying type prediction vector with historical behavior.

[0064] Therefore, it can be understood that the crying classification prediction device determines the historical behavior interval time based on the historical behavior time of the historical behavior information. Further, the crying classification prediction device determines the probability parameters of the historical behavior cause based on the historical behavior interval time, and determines the adjustment vector based on the probability parameters.

[0065] It should be noted that the historical behavior adjustment module is a module that supplements the crying type prediction vector output by the classification prediction module with historical behavior. Therefore, the formula of the historical behavior adjustment module is described as:

[0066] h(M)=N

[0067] Wherein, M is the input data collected by the data collection module, N is the adjustment vector, and h(·) is the mapping algorithm used by the historical behavior adjustment module. The mapping algorithm is used to map the historical behavior information into the adjustment vector. The mapping algorithm can be replaced according to the needs.

[0068] It should be noted that the historical behavior information adjustment and the historical behavior recording function complement and promote each other, that is, the historical behavior information adjustment depends on the data provided by the historical behavior recording function, and the historical behavior information adjustment will also require the historical behavior recording function to be expanded and improved. In addition, the continuously expanded historical behavior recording function can provide users with more abundant functions, such as daily reports on baby conditions and baby health analysis, etc. These extended functions can enhance product value and increase user stickiness.

[0069] Further, refer to Figure 2It can be seen that the crying classification prediction device includes a feature fusion module, which is used to fuse the crying type prediction vector output by the classification prediction module with the adjustment vector output by the historical behavior adjustment module. Therefore, the crying classification prediction device fuses the crying type prediction vector with the adjustment vector to obtain a target prediction result, wherein the target prediction result includes the crying type and the probability value of the crying type.

[0070] It should be noted that the feature fusion module is based on the idea of ​​the late fusion method, which fuses the cry type prediction vector output by the classification prediction module with the adjustment vector output by the historical behavior adjustment module to obtain the final output result of the entire system. Therefore, the formula of the feature fusion module is described as:

[0071] g(Y,N)=Z

[0072] Wherein, Y is the cry type prediction vector, N is the adjustment vector, g(·) is the fusion algorithm used by the feature fusion module, and the fusion algorithm is used to use the adjustment vector N to supplement or correct the cry type prediction vector Y and obtain the final target prediction result Z. The fusion algorithm can be replaced according to needs.

[0073] It should be further explained that the embodiment of the present invention can also implement the crying classification prediction process without historical behavior information. Therefore, the feature fusion module should also meet the following conditions: when the user uses the crying classification prediction device through the user terminal for the first time, the behavior recording module does not collect historical behavior information, then the adjustment vector output by the historical behavior adjustment module is N=0, and the formula of the feature fusion module is described as:

[0074] g(Y,0)=Y

[0075] In other words, without using historical behavior information, the feature fusion module cannot affect the output of the classification prediction module, and the final output of the system is equal to the output of the classification prediction module.

[0076] Further, refer to Figure 2 It can be seen that the crying classification prediction device includes an output module, which is used to output the target prediction result obtained after fusion and send it to the user terminal. Therefore, the crying classification prediction device outputs the target prediction result and sends it to the user terminal. Further, after the user receives the target prediction result through the user terminal, the target user's crying root cause is determined according to the target prediction result, and corresponding solutions are taken according to the target user's crying root cause. In one embodiment, if the target user's crying root cause is hunger, then the user takes the corresponding solution of feeding the target user. In another embodiment, if the target user's crying root cause is pain caused by stumbling, then the user takes the corresponding solution of inspecting and treating the target user's wound.

[0077] The crying classification prediction method provided by the embodiment of the present invention takes the audio sequence and historical behavior information as independent modal data respectively, obtains the crying type prediction vector by predicting the crying type of the audio sequence, obtains the adjustment vector by the adjustment mode of the historical behavior, and then predicts the target prediction result by fusing the outputs of each modality to achieve the adjustment of the crying type prediction vector. The accuracy of crying classification prediction can be improved, and the calculation difficulty can be reduced through a lightweight multi-modal structure, thereby improving the efficiency of crying classification prediction.

[0078] Furthermore, the predicting the crying type based on the audio sequence to obtain the crying type prediction vector includes:

[0079] Performing audio preprocessing on the audio sequence to obtain at least one audio segment feature;

[0080] Inputting each of the audio segment features into a crying type prediction model to obtain at least one crying type vector of the segment output by the crying type prediction model; the crying type prediction model is obtained by training an audio neural network;

[0081] Based on the cry type vectors of each segment, a cry type prediction vector is determined.

[0082] Specifically, the crying classification prediction device performs audio preprocessing on the audio sequence, that is, converts the audio sequence from a time domain signal into an easy-to-process frequency domain signal, and obtains at least one audio segment feature, wherein the audio segment feature refers to a spectrogram feature, and the spectrogram feature includes but is not limited to a Mel Frequency Cepstrum Coefficient feature (MFCC) and a Linear Prediction Cepstrum Coefficient feature (LPCC).

[0083] Furthermore, the crying classification prediction device inputs the features of each audio segment into the crying type prediction model to obtain at least one segment crying type vector output by the crying type prediction model and the probability value corresponding to the segment crying type vector, wherein the crying type prediction model is obtained by training the audio neural network, and the audio neural network can select the neural network model according to the actual situation. In one embodiment, the audio neural network is a ResNet18 network model. ResNet-18 is a deep residual network architecture that introduces the concept of residual connection. It can directly skip the output of some layers, so that the network can better train deep structures.

[0084] It should be noted that during the training process, the crying type prediction model uses at least one crying type as sample data and inputs it into the audio neural network in a fixed order for training.

[0085] In one embodiment, there are three types of crying: "non-crying", "non-hunger crying" and "hunger crying". In the model training process, the sample audio data and the crying type vector label data are input as sample data into the audio neural network for training to obtain a crying type prediction model, wherein the crying type vector label data includes three components of three crying categories: "non-crying", "non-hunger crying" and "hunger crying", and the probability value of the "non-crying" type is c1, the probability value of the "non-hunger crying" type is c2, and the probability value of the "hunger crying" type is c3. Therefore, the crying type vector label data can be represented as a 1×3 vector of Y=[c1, c2, c3]. In the process of crying classification prediction, the crying classification prediction device inputs the audio segment feature into the crying type prediction model to obtain the segment crying type vector corresponding to the audio segment feature, and each component of the segment crying type vector represents the probability value of the "non-crying" type, the probability value of the "non-hunger crying" type and the probability value of the "hunger crying" type.

[0086] Furthermore, the crying classification prediction device determines a crying type prediction vector based on the crying type vector of each segment.

[0087] The embodiment of the present invention obtains at least one audio segment feature by performing audio preprocessing on an audio sequence, and further inputs each audio segment feature into a crying type prediction model to obtain at least one segment crying type vector output by the crying type prediction model, and determines the crying type prediction vector based on each segment crying type vector. By taking the classification prediction module as the main body of the algorithm and making it independent of the historical behavior adjustment module, the algorithm can be freely replaced without affecting the prediction result. It can select a suitable model or algorithm based on the actual situation, and can also use new models and algorithms proposed in cutting-edge research in a timely manner. It has high interchangeability and improves the flexibility of the system. At the same time, by obtaining simpler modal information, it avoids the use of special or professional equipment to reduce the difficulty of product deployment and promotion of the technical solution.

[0088] Furthermore, the performing audio preprocessing on the audio sequence to obtain at least one audio segment feature includes:

[0089] Segmenting the audio sequence to obtain at least one audio segment;

[0090] Perform spectrum conversion on each of the audio segments to obtain at least one audio segment feature.

[0091] Specifically, the crying classification prediction device determines the length of the audio sequence based on the acquired audio sequence. Further, the crying classification prediction device divides the audio sequence length according to the preset length, that is, divides the audio sequence into parts to obtain at least one audio segment, wherein the preset length is set according to actual conditions.

[0092] In one embodiment, if the duration of the audio sequence is 25 seconds, the crying classification prediction device divides the audio sequence into 5-second segments to obtain 5 audio segments.

[0093] Furthermore, the crying sound classification prediction device performs spectrum conversion on each audio segment to obtain at least one audio segment feature, that is, obtains at least one spectrogram feature. It should be noted that each audio segment will obtain an audio segment feature after the spectrum conversion.

[0094] In one embodiment, the spectrum conversion includes short-time Fourier transform and Mel scale transform. Short-time Fourier transform is used to convert the signal from the time domain to the frequency domain. Mel scale transform is a nonlinear scale that converts the frequency into the pitch perceived by the human ear, and is usually used to convert the frequency of the audio signal into a Mel spectrogram. Therefore, the crying classification prediction device performs short-time Fourier transform and Mel scale transform on each audio segment to obtain a Mel spectrogram after the short-time Fourier transform and Mel scale transform of each audio segment.

[0095] The embodiment of the present invention divides the audio sequence into segments to obtain at least one audio segment, performs spectrum conversion on each audio segment to obtain at least one audio segment feature, and obtains multiple audio segment features by preprocessing the audio sequence, which helps to reduce computational complexity, facilitates prediction processing by a crying type prediction model, and improves the generalization performance and accuracy of the model.

[0096] Furthermore, the determining of the crying type prediction vector based on the crying type vectors of each segment includes:

[0097] The cry type vectors of each segment are averaged to obtain the cry type prediction vector.

[0098] Specifically, the crying classification prediction device performs mean calculation on the crying type vectors of each segment and the probability values ​​corresponding to the crying type vectors of each segment output by the crying type prediction model to obtain the crying type prediction vector, wherein the mean calculation formula is:

[0099]

[0100] Among them, n represents the crying type vector of n segments, Y i Represents the crying type vector of the i-th segment.

[0101] The embodiment of the present invention obtains a crying type prediction vector by performing mean calculation on the crying type vectors of each segment. The mean calculation can compress vectors of multiple dimensions into a vector of one dimension, thereby reducing the difference influence of the crying type vectors of each segment, improving the robustness of the model, and thus reducing the complexity of the fusion calculation performed by the feature fusion module.

[0102] Further, determining the adjustment vector of the target user based on the historical behavior information includes:

[0103] Determine the historical behavior interval time based on the historical behavior time of the historical behavior information;

[0104] Determining a probability parameter of a cause of the historical behavior based on the historical behavior interval;

[0105] Based on the probability parameter, the adjustment vector is determined.

[0106] Specifically, the crying classification prediction device determines the historical behavior time when the historical behavior information occurred and the current time based on the acquired historical behavior information. Furthermore, the crying classification prediction device subtracts the current time from the historical behavior time based on the historical behavior time to obtain the historical behavior interval time.

[0107] Furthermore, the crying classification prediction device determines the probability parameters of the historical behavior causes based on the historical behavior interval time, wherein the calculation of the probability parameters of the historical behavior causes is determined according to actual conditions.

[0108] Furthermore, the crying classification prediction device determines the adjustment vector based on the probability parameter.

[0109] In one embodiment, taking the example of determining whether the target user is crying due to hunger, the historical behavior information is the feeding behavior record. The crying classification prediction device determines the last feeding time point as t based on the feeding behavior record, and determines the current time point t'. Further, the crying classification prediction device subtracts the current time point from the last feeding time point to obtain the historical behavior interval time as △t=t'-t, and converts the historical behavior interval time into hours.

[0110] Furthermore, the crying classification prediction device determines that the feeding interval of a general baby is 3 hours based on the medical observation of the target user's recommended feeding interval, and constructs a probability parameter λ∈(0,2) of baby hunger based on the medical observation. Then the probability parameter is calculated as follows:

[0111] λ=1-tanh(3-△t)

[0112] Among them, tanh(·) is a hyperbolic tangent function. When △t approaches positive infinity, the value of the function approaches 0, and when △t approaches negative infinity, the value of the function approaches 2. Therefore, the range of the function is (0,2). Therefore, the larger the historical behavior interval, the larger the probability parameter value, and the more likely the target user is crying because of hunger; the smaller the historical behavior interval, the smaller the probability parameter value, and the less likely the target user is crying because of hunger. In particular, when there is no feeding record, that is, when the last feeding time point t does not exist, λ=1.

[0113] Furthermore, the crying classification prediction device determines the adjustment vector as N=[1,2-λ,λ] based on the calculated probability parameters.

[0114] The embodiment of the present invention determines the historical behavior interval time based on the historical behavior time of the historical behavior information, further determines the probability parameters of the historical behavior causes based on the historical behavior interval time, and determines the adjustment vector based on the probability parameters to achieve historical behavior supplementation of the crying type prediction vector. By using the historical behavior adjustment module as an independent algorithm module, data mining is performed on the historical behavior information, and it is independent of the classification prediction module, so that the algorithm can be freely replaced without affecting the prediction results. It has high interchangeability and strong versatility, which improves the flexibility of the system. At the same time, historical behavior information is optional supplementary information, and even the lack of historical behavior information will not have a negative impact on the prediction results, thereby achieving an adjustment design that allows the missing of some modal information and effectively improving the algorithm performance.

[0115] Furthermore, the step of fusing the cry type prediction vector and the adjustment vector to obtain a target prediction result includes:

[0116] Performing matrix transposition processing on the adjustment vector to obtain an adjustment transposed vector;

[0117] Performing a dot product calculation on the crying type prediction vector and the adjusted transposed vector to obtain a target prediction vector;

[0118] Based on the target prediction vector, determining a target probability value for each crying type in the target prediction vector;

[0119] A target prediction result is determined based on the target probability value of each cry type in the target prediction vector.

[0120] Specifically, the crying classification prediction device performs matrix transposition processing on the adjustment vector to obtain the adjustment transposed vector. Further, the crying classification prediction device performs point product calculation on the crying type prediction vector and the adjustment transposed vector to obtain the target prediction vector.

[0121] In one embodiment, the crying type prediction vector is Y=[c1, c2, c3], and the adjustment vector is N=[1, 2-λ, λ]. Then the crying classification prediction device performs matrix transposition processing on the adjustment vector to obtain the adjustment transposed vector:

[0122]

[0123] Then the target prediction vector obtained by the crying classification prediction device by performing a dot multiplication calculation on the crying type prediction vector and the adjustment transposed vector is:

[0124]

[0125] Furthermore, the crying classification prediction device converts the target prediction vector into a probability value based on the softmax function, and determines the target probability value of each crying type in the target prediction vector. Further, the crying classification prediction device determines the target prediction result based on the target probability value of each crying type in the target prediction vector, wherein the target prediction result includes the crying type and the probability value of the crying type.

[0126] The embodiment of the present invention performs dot product calculation on the adjusted transposed vector obtained by matrix transposing the crying type prediction vector and the adjustment vector to obtain a target prediction vector, and determines the target probability value of each crying type in the target prediction vector based on the target prediction vector, and determines the target prediction result based on the target probability value of each crying type in the target prediction vector. The target prediction result is predicted by fusing the outputs of each modality to achieve the adjustment of the crying type prediction vector, which can improve the accuracy of crying classification prediction, and can also improve the efficiency of crying classification prediction through a lightweight multi-modal structure. At the same time, based on a simple modal fusion mechanism, the algorithm calculation complexity can be reduced and the algorithm calculation overhead can be reduced.

[0127] Further, the determining of the target prediction result based on the target probability value of each crying type in the target prediction vector includes:

[0128] Determining a maximum target probability value based on the target probability values ​​of each cry type in the target prediction vector;

[0129] Based on the maximum target probability value, a target prediction result is determined.

[0130] Specifically, the crying classification prediction device compares the target probability values ​​of each crying type in the target prediction vector based on the target probability values ​​of each crying type in the target prediction vector to obtain a comparison result. Further, the crying classification prediction device determines the maximum target probability value based on the comparison result.

[0131] Furthermore, the crying classification prediction device determines the crying type corresponding to the maximum target probability value based on the maximum target probability value. Furthermore, the crying classification prediction device associates the maximum target probability value with the crying type corresponding to the determined maximum target probability value to obtain a target prediction result, and outputs the target prediction result, so that the user can receive the target prediction result output by the crying classification prediction device through the user terminal, and determine the root cause of the target user's crying based on the target prediction result, and then take corresponding solutions according to the root cause of the target user's crying.

[0132] The embodiment of the present invention determines the maximum target probability value based on the target probability value of each crying type in the target prediction vector, and determines the target prediction result based on the maximum target probability value. The target prediction result determined by the maximum probability value can improve the prediction accuracy and enhance the interpretability of the result, thereby realizing the adjustment of the crying type prediction vector, improving the accuracy of crying classification prediction, and improving the efficiency of crying classification prediction through a lightweight multimodal structure.

[0133] Further, refer to Figure 3 , Figure 3 : is a flow chart of the overall solution of the crying sound classification prediction method provided by an embodiment of the present invention. Therefore, the overall process of the crying sound classification prediction method provided by the present invention can be understood as follows:

[0134] The crying sound classification prediction device obtains the audio data and historical behavior information of the target user. Further, the crying sound classification prediction device performs sequence conversion on the obtained audio data of the target user to obtain the audio sequence of the target user.

[0135] Furthermore, the crying sound classification prediction device determines the duration of the audio sequence based on the acquired audio sequence. Furthermore, the crying sound classification prediction device divides the duration of the audio sequence according to the preset duration, that is, divides the audio sequence into segments to obtain at least one audio segment.

[0136] Furthermore, the crying sound classification prediction device performs spectrum conversion on each audio segment to obtain at least one audio segment feature, that is, obtains at least one spectrum graph feature.

[0137] Furthermore, the crying classification prediction device inputs the features of each audio segment into a crying type prediction model to obtain at least one segment crying type vector output by the crying type prediction model and a probability value corresponding to the segment crying type vector.

[0138] Furthermore, the crying classification prediction device performs mean calculation on each segment crying type vector output by the crying type prediction model and the probability value corresponding to the segment crying type vector to obtain the crying type prediction vector.

[0139] Furthermore, the crying classification prediction device determines the historical behavior time when the historical behavior information occurred and the current time based on the acquired historical behavior information. Furthermore, the crying classification prediction device subtracts the current time from the historical behavior time based on the historical behavior time to obtain the historical behavior interval time.

[0140] Further, the crying sound classification prediction device determines the probability parameters of the historical behavior causes based on the historical behavior interval time. Further, the crying sound classification prediction device determines the adjustment vector based on the probability parameters.

[0141] Further, the crying classification prediction device performs matrix transposition processing on the adjustment vector to obtain the adjustment transposed vector. Further, the crying classification prediction device performs dot product calculation on the crying type prediction vector and the adjustment transposed vector to obtain the target prediction vector.

[0142] Furthermore, the crying classification prediction device converts the target prediction vector into a probability value based on the softmax function, and determines the target probability value of each crying type in the target prediction vector.

[0143] Furthermore, the crying classification prediction device compares the target probability values ​​of each crying type in the target prediction vector based on the target probability values ​​of each crying type in the target prediction vector to obtain a comparison result. Further, the crying classification prediction device determines the maximum target probability value based on the comparison result.

[0144] Furthermore, the crying classification prediction device determines the crying type corresponding to the maximum target probability value based on the maximum target probability value. Furthermore, the crying classification prediction device associates the maximum target probability value with the crying type corresponding to the determined maximum target probability value to obtain a target prediction result.

[0145] Furthermore, the crying classification prediction device outputs the target prediction result and sends it to the user terminal. Furthermore, after the user receives the target prediction result through the user terminal, the root cause of the target user's crying is determined according to the target prediction result, and corresponding solutions are taken according to the root cause of the target user's crying.

[0146] Furthermore, the present invention also provides a crying sound classification prediction device.

[0147] Reference Figure 4 , Figure 4 This is a schematic diagram of the functional modules of an embodiment of a crying sound classification prediction device of the present invention.

[0148] The crying sound classification prediction device comprises:

[0149] An acquisition module 410 is used to acquire the audio sequence and historical behavior information of the target user;

[0150] A prediction module 420, configured to predict the crying type based on the audio sequence to obtain a crying type prediction vector;

[0151] An adjustment module 440 is used to determine an adjustment vector of the target user based on the historical behavior information; the adjustment vector is used to supplement the crying type prediction vector with historical behavior;

[0152] The fusion module 450 is used to fuse the cry type prediction vector and the adjustment vector to obtain a target prediction result.

[0153] The crying classification prediction device provided in the embodiment of the present invention takes the audio sequence and historical behavior information as independent modal data respectively, obtains the crying type prediction vector by predicting the crying type of the audio sequence, obtains the adjustment vector by the adjustment mode of the historical behavior, and then predicts the target prediction result by fusing the outputs of each modality to achieve the adjustment of the crying type prediction vector. This can improve the accuracy of crying classification prediction, and can also reduce the calculation difficulty and improve the efficiency of crying classification prediction through a lightweight multi-modal structure.

[0154] In one embodiment, the prediction module 420 is further configured to:

[0155] Performing audio preprocessing on the audio sequence to obtain at least one audio segment feature;

[0156] Inputting each of the audio segment features into a crying type prediction model to obtain at least one crying type vector of the segment output by the crying type prediction model; the crying type prediction model is obtained by training an audio neural network;

[0157] Based on the cry type vectors of each segment, a cry type prediction vector is determined.

[0158] In one embodiment, the prediction module 420 is further configured to:

[0159] Segmenting the audio sequence to obtain at least one audio segment;

[0160] Perform spectrum conversion on each of the audio segments to obtain at least one audio segment feature.

[0161] In one embodiment, the prediction module 420 is further configured to:

[0162] The cry type vectors of each segment are averaged to obtain the cry type prediction vector.

[0163] In one embodiment, the adjustment module 440 is further configured to:

[0164] Determine the historical behavior interval time based on the historical behavior time of the historical behavior information;

[0165] Determining a probability parameter of a cause of the historical behavior based on the historical behavior interval;

[0166] Based on the probability parameter, the adjustment vector is determined.

[0167] In one embodiment, the fusion module 450 is further configured to:

[0168] Performing matrix transposition processing on the adjustment vector to obtain an adjustment transposed vector;

[0169] Performing a dot product calculation on the crying type prediction vector and the adjusted transposed vector to obtain a target prediction vector;

[0170] Based on the target prediction vector, determining a target probability value for each crying type in the target prediction vector;

[0171] A target prediction result is determined based on the target probability value of each cry type in the target prediction vector.

[0172] In one embodiment, the fusion module 450 is further configured to:

[0173] Determining a maximum target probability value based on the target probability values ​​of each cry type in the target prediction vector;

[0174] Based on the maximum target probability value, a target prediction result is determined.

[0175] Figure 5 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 5 As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 540 and a communication bus 550, wherein the processor 510, the communication interface 520 and the memory 540 communicate with each other through the communication bus 550. The processor 510 may call a computer program in the memory 540 to execute the steps of the crying classification prediction method, for example, including:

[0176] Obtain the target user's audio sequence and historical behavior information;

[0177] Predicting the crying type based on the audio sequence to obtain a crying type prediction vector;

[0178] Based on the historical behavior information, determining an adjustment vector of the target user; the adjustment vector is used to supplement the crying type prediction vector with historical behavior;

[0179] The cry type prediction vector and the adjustment vector are fused to obtain a target prediction result.

[0180] In addition, the logic instructions in the above-mentioned memory 540 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0181] On the other hand, an embodiment of the present invention further provides a medium, which is a computer-readable storage medium, and the computer-readable storage medium stores a computer program, and the computer program is used to enable a processor to execute the steps of the methods provided in the above embodiments, for example, including:

[0182] Obtain the target user's audio sequence and historical behavior information;

[0183] Predicting the crying type based on the audio sequence to obtain a crying type prediction vector;

[0184] Based on the historical behavior information, determining an adjustment vector of the target user; the adjustment vector is used to supplement the crying type prediction vector with historical behavior;

[0185] The cry type prediction vector and the adjustment vector are fused to obtain a target prediction result.

[0186] The computer-readable storage medium can be any available medium or data storage device that can be accessed by the processor, including but not limited to magnetic storage (such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc.), optical storage (such as CD, DVD, BD, HVD, etc.), and semiconductor storage (such as ROM, EPROM, EEPROM, non-volatile memory (NANDFLASH), solid-state drive (SSD)), etc.

[0187] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0188] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0189] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A crying sound classification prediction method, characterized in that: include: Obtain the target user's audio sequence and historical behavior information; Predicting the crying type based on the audio sequence to obtain a crying type prediction vector; Based on the historical behavior information, determining an adjustment vector of the target user; the adjustment vector is used to supplement the crying type prediction vector with historical behavior; The cry type prediction vector and the adjustment vector are fused to obtain a target prediction result.

2. The crying sound classification prediction method according to claim 1, characterized in that: The determining the adjustment vector of the target user based on the historical behavior information includes: Determine the historical behavior interval time based on the historical behavior time of the historical behavior information; Determining a probability parameter of a cause of the historical behavior based on the historical behavior interval; Based on the probability parameter, the adjustment vector is determined.

3. The crying sound classification prediction method according to claim 1, characterized in that: The step of fusing the cry type prediction vector and the adjustment vector to obtain a target prediction result includes: Performing matrix transposition processing on the adjustment vector to obtain an adjustment transposed vector; Performing a dot product calculation on the crying type prediction vector and the adjusted transposed vector to obtain a target prediction vector; Based on the target prediction vector, determining a target probability value for each crying type in the target prediction vector; A target prediction result is determined based on the target probability value of each cry type in the target prediction vector.

4. The crying sound classification prediction method according to claim 3, characterized in that: The determining of the target prediction result based on the target probability value of each crying type in the target prediction vector includes: Determining a maximum target probability value based on the target probability values ​​of each cry type in the target prediction vector; Based on the maximum target probability value, a target prediction result is determined.

5. The crying sound classification prediction method according to claim 1, characterized in that: The step of predicting the crying type based on the audio sequence to obtain a crying type prediction vector includes: Performing audio preprocessing on the audio sequence to obtain at least one audio segment feature; Inputting each of the audio segment features into a crying type prediction model to obtain at least one crying type vector of the segment output by the crying type prediction model; the crying type prediction model is obtained by training an audio neural network; Based on the cry type vectors of each segment, a cry type prediction vector is determined.

6. The crying sound classification prediction method according to claim 5, characterized in that: The performing audio preprocessing on the audio sequence to obtain at least one audio segment feature includes: Segmenting the audio sequence to obtain at least one audio segment; Perform spectrum conversion on each of the audio segments to obtain at least one audio segment feature.

7. The crying sound classification prediction method according to claim 5, characterized in that: The step of determining a crying type prediction vector based on the crying type vectors of each segment includes: The cry type vectors of each segment are averaged to obtain the cry type prediction vector.

8. A crying sound classification prediction device, characterized in that: include: The acquisition module is used to obtain the audio sequence and historical behavior information of the target user; A prediction module, used for predicting the crying type based on the audio sequence to obtain a crying type prediction vector; An adjustment module, used for determining an adjustment vector of the target user based on the historical behavior information; the adjustment vector is used for supplementing the crying type prediction vector with historical behavior; A fusion module is used to fuse the cry type prediction vector and the adjustment vector to obtain a target prediction result.

9. An electronic device comprising a processor and a memory storing a computer program, characterized in that: When the processor executes the computer program, the crying sound classification prediction method according to any one of claims 1 to 7 is implemented.

10. A storage medium, the storage medium being a computer-readable storage medium, comprising a computer program, characterized in that: When the computer program is executed by a processor, the crying sound classification prediction method according to any one of claims 1 to 7 is implemented.