Sound data processing method and apparatus

By converting sound data into latent space fusion vectors for clustering, the problem of inaccurate sound feature extraction and analysis is solved, resulting in more natural and stable clustering results.

CN120636414BActive Publication Date: 2026-02-17BEIJING XIYU JIZHI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511004663.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2026-02-17
Estimated Expiration
2045-07-21

AI Technical Summary

Technical Problem

Existing sound data processing technologies cannot guarantee the accuracy of clustering results, leading to inaccurate sound feature extraction and analysis.

Method used

By identifying the audio dataset and converting each audio data point into a latent space fusion vector based on the sound data processing model, latent space clustering is performed to obtain sound clustering information.

Benefits of technology

It achieves accurate extraction and analysis of sound features, and the clustering results are more natural and stable, avoiding inaccuracies in the clustering results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636414B_ABST
    Figure CN120636414B_ABST
Patent Text Reader

Abstract

The application discloses a sound data processing method and device. The method comprises the following steps: determining an audio data set, wherein the audio data set comprises at least one piece of sound data; processing the audio data set based on a sound data processing model to obtain at least one piece of sound clustering information; each piece of sound clustering information comprises acoustic characteristic information of a same object; wherein the sound data processing model is used for converting each piece of sound data into an implicit space fusion vector, and performing clustering processing based on the implicit space fusion vector to obtain the sound clustering information. The technical scheme of the application realizes accurate extraction and analysis of sound characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of sound recognition, and in particular to a sound data processing method and device. BACKGROUND

[0002] Sound data processing technology has deeply penetrated into many fields such as daily life, industrial production and public service due to its ability of analyzing, understanding and transforming audio signals, and its application scenarios cover not only basic human-computer interaction but also complex security verification and emotional computing.

[0003] At present, sound data processing technology usually collects different voiceprint features in audio signals, and then directly performs clustering analysis on the voiceprint features to obtain clustering results for distinguishing voice information of different objects. However, the above-mentioned simple and direct clustering method cannot guarantee that the clustering results can reflect the voice information of the same object, that is, there is still a problem of inaccurate clustering results, so it is very important to accurately extract and analyze sound features. SUMMARY

[0004] The present application provides a sound data processing method and device to accurately extract and analyze sound features.

[0005] According to an aspect of the present application, a sound data processing method is provided, which comprises:

[0006] determining an audio data set containing at least one piece of sound data;

[0007] processing the audio data set based on a sound data processing model to obtain at least one piece of sound clustering information; each piece of sound clustering information contains acoustic feature information of the same object; wherein the sound data processing model is used to convert each piece of sound data into an embedding fusion vector, and perform clustering processing based on the embedding fusion vector to obtain sound clustering information.

[0008] According to another aspect of the present application, a sound data processing device is provided, which comprises:

[0009] a data set determination module for determining an audio data set containing at least one piece of sound data;

[0010] a processing module for processing the audio data set based on a sound data processing model to obtain at least one piece of sound clustering information; each piece of sound clustering information contains acoustic feature information of the same object; wherein the sound data processing model is used to convert each piece of sound data into an embedding fusion vector, and perform clustering processing based on the embedding fusion vector to obtain sound clustering information.

[0011] According to another aspect of the present application, there is provided an electronic device comprising:

[0012] at least one processor; and

[0013] a memory communicatively connected to the at least one processor; wherein

[0014] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the sound data processing method according to any one of the embodiments of the present application.

[0015] According to another aspect of the present application, there is provided a computer readable storage medium storing computer instructions for enabling a processor to implement the sound data processing method according to any one of the embodiments of the present application when executed by the processor.

[0016] The technical solution of the embodiments of the present application determines at least one sound data audio data set, and processes the audio data set based on a sound data processing model to obtain at least one sound clustering information; each sound clustering information contains acoustic feature information of the same object, thereby realizing extraction of acoustic feature information of each object; and the sound data processing model of the present application is used to convert each piece of sound data into a latent space fusion vector, and perform clustering processing based on the latent space fusion vector to obtain sound clustering information. The effect of fusion in the latent space is more natural and stable than fusion in the explicit space, thereby more ensuring accurate extraction and analysis of sound features.

[0017] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0019] Figure 1 is a flowchart of a sound data processing method according to an embodiment of the present application;

[0020] Figure 2 is a flowchart of another sound data processing method according to an embodiment of the present application;

[0021] Figure 3 is a flow chart of another sound data processing method according to an embodiment of the present application;

[0022] Figure 4 is a structural schematic diagram of a sound data processing device according to an embodiment of the present application;

[0023] Figure 5 is a structural schematic diagram of an electronic device for implementing the sound data processing method according to an embodiment of the present application. DETAILED DESCRIPTION

[0024] In order to make the personnel in the art better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should belong to the scope of protection of the present application.

[0025] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0026] Embodiment one

[0027] Figure 1 A flow chart of a sound data processing method according to an embodiment of the present application is provided, the embodiment can be applicable to the case of processing sound data, the method can be executed by a sound data processing device, the sound data processing device can be realized in the form of hardware and / or software, and the sound data processing device can be configured in any electronic device with network communication function. As shown in the figure, the sound data processing method of the present application includes the following processes: Figure 1

[0028] S110, determining an audio data set, the audio data set containing at least one piece of sound data.

[0029] ​Each piece of sound data corresponds to an object, ensuring the purity of each piece of sound data and avoiding multiple sounds being mixed together, which leads to low precision in sound feature extraction.

[0030] Specifically, determining the audio data set can include the following steps A1-A5:

[0031] Step A1, obtaining a plurality of pieces of sound data to be processed.

[0032] Step A2, determining the number of sound-emitting objects included in each piece of sound data to be processed.

[0033] Step A3, if the sound data to be processed contains the sound data of only one object, the sound data to be processed is stored in the audio data set; if the sound data to be processed contains the sound data of at least two objects, it is determined whether the sound data of each object in the sound data to be processed can be processed by segmentation.

[0034] Step A4, if the sound data of each object in the sound data to be processed can be separated by audio segmentation, the sound data of each object is processed by separation, and the sound data processed by segmentation is stored in the audio data set.

[0035] Specifically, after processing the sound data of each object by segmentation, the sound data processed by segmentation is determined as first sound data, which can be directly stored in the audio data set, or the first sound data belonging to the same object is spliced to obtain second sound data, and the second sound data is stored in the audio data set.

[0036] Step A5, if the sound data of each object in the sound data to be processed cannot be processed by segmentation, the sound data to be processed is rejected.

[0037] For example, the sound data to be processed contains the sound data of object A and the sound data of object B, and the sound data to be processed is the conversation content of object A and object B, that is, object A and object B do not speak at the same time. Therefore, the sound data of object A and the sound data of object B need to be separated from the sound data to be processed, and after separation, they can be directly stored in the audio data set, or the segmented data belonging to object A can be spliced together and stored in the audio data set, and the segmented data belonging to object B can be spliced together and stored in the audio data set.

[0038] S120, processing the audio data set based on a sound data processing model to obtain at least one sound clustering information; each sound clustering information contains acoustic feature information of the same object; wherein the sound data processing model is used to convert each piece of sound data into a hidden space fusion vector, and perform clustering processing based on the hidden space fusion vector to obtain sound clustering information.

[0039] The acoustic feature information can be understood as a stable and unique voiceprint feature that can reflect the unique voiceprint of the object, including but not limited to frequency, tone, speech rate, tone, and other acoustic information. The latent space fusion vector can be understood as a fusion vector obtained by mapping the sound features in each piece of sound data to the latent space and then performing data fusion on the at least two latent space vectors in the latent space.

[0040] Specifically, the sound data processing model extracts sound features from each piece of sound data in the audio data set to obtain acoustic feature information, converts the acoustic feature information into a latent space fusion vector, and performs clustering processing based on the latent space fusion vector to obtain sound clustering information.

[0041] In the embodiments of the present application, the clustering processing based on the latent space fusion vector to obtain the sound clustering information includes: determining the latent space distance between the latent space fusion vectors of each piece of sound data; and determining the clustering result of each piece of sound data based on the comparison result of the latent space distance and the maximum clustering distance threshold.

[0042] The calculation method of the space distance in the latent space is the same as that of the space distance in the explicit space, for example, cosine similarity, Euclidean distance, etc.

[0043] Specifically, the maximum clustering distance threshold can be understood as the maximum distance between the latent space fusion vectors belonging to the same object, and the determination of the clustering result of each piece of sound data based on the comparison result of the latent space distance and the maximum clustering distance threshold can include: when the latent space distance is less than or equal to the maximum clustering distance threshold, the acoustic feature information corresponding to the corresponding latent space fusion vector belongs to the same object; and when the latent space distance is greater than the maximum clustering distance threshold, the acoustic feature information corresponding to the corresponding latent space fusion vector does not belong to the same object.

[0044] The technical scheme of the present embodiment determines the latent space distance between the latent space fusion vectors of each piece of sound data. Because the dimension of the latent vector is lower, performing fusion and space distance determination in the latent space can greatly improve the calculation efficiency. Further, based on the comparison result of the latent space distance and the maximum clustering distance threshold, the clustering result of each piece of sound data is determined. Compared with the sound vector fusion in the explicit space, the feature density in the latent space is greater, and performing vector fusion in the latent space can better reflect the deep features of the sound, such as tone and other difficult-to-quantify features. Therefore, clustering processing of the latent vector in the latent space can be more natural and stable, and can better reflect the sound features of different objects.

[0045] Optionally, the clustering processing in the sound data processing model is completed in a clustering layer in the sound data processing model. The sound data processing model can be trained in a contrastive learning manner, and the specific process can be: constructing a contrastive learning database, the contrastive learning database can include positive example data and negative example data; wherein the positive example data is sound data of the same object, and the negative example data is sound data of different objects; in order to enhance the contrastive learning effect, the negative example data includes relatively obvious non-same object sound data, and also includes negative example data with certain similarity, the negative example data with certain similarity means that the object timbre of the negative example data is similar to at least one dimension of the acoustic feature of the target object, but the object of the negative example data is not the same object as the target object, so as to strengthen the learning of the model for similar acoustic features.

[0046] Further, the data in the contrastive learning database is input into the sound data processing model to be trained to obtain predicted data, a first loss value of the predicted data with respect to the positive example data and the negative example data in the contrastive learning database is determined, all the first loss values are averaged to obtain a second loss value, and the parameters of the sound data processing model to be trained are optimized according to the second loss value until convergence, thereby obtaining the sound data processing model.

[0047] Loss value L q The calculation process can be obtained by using a contrastive learning loss (InfoNCE Loss), which can be represented by the following formula:

[0048]

[0049] Wherein, τ is a temperature coefficient for controlling the discrimination degree of the model to the negative sample. q is a feature encoding, q and the positive example data k + are positive example data pairs, q and the negative example data k i are positive example data pairs. q·k is the logits output by the model, which refers to the original score of the model output without softmax normalization. These scores represent the prediction confidence of the model for each object. Specifically, logits are the score values directly output by the model after processing the input data. These scores have not been converted into probability values, so they can directly reflect the preference degree of the model for different objects. Logits specifically refer to the unnormalized scores output by the model during contrastive learning. They are the original prediction results obtained by the model after processing the input data. These logits are combined with the temperature coefficient when calculating InfoNCEloss, to adjust the prediction distribution of the model for different objects.

[0050] The technical scheme of the embodiment of the present application determines an audio data set containing at least one piece of sound data, processes the audio data set based on a sound data processing model, and obtains at least one piece of sound clustering information; each piece of sound clustering information contains acoustic characteristic information of the same object, thereby realizing extraction of acoustic characteristic information of each object; and the sound data processing model of the present application is used to convert each piece of sound data into a latent space fusion vector, and perform clustering processing based on the latent space fusion vector to obtain sound clustering information. The effect of fusion in the latent space is more natural and stable than fusion in the explicit space, thereby more ensuring accurate extraction and analysis of sound characteristics.

[0051] Embodiment two

[0052] Figure 2 The flowchart of another sound data processing method provided by the embodiment of the present application, the technical scheme of the present embodiment is further optimized for the process of S120 in the foregoing embodiment on the basis of the foregoing embodiment, and the present embodiment can be combined with each optional scheme in one or more of the foregoing embodiments. As shown in Figure 2 The sound data processing method of the present application includes the following processes:

[0053] S210, determining an audio data set and a sound data processing model; the audio data set contains at least one piece of sound data; the sound data processing model includes an acoustic characteristic extraction layer, which includes a first acoustic characteristic extraction layer and a second acoustic characteristic extraction layer.

[0054] S220, performing sound feature extraction on each piece of sound data in the audio data set based on the first acoustic characteristic extraction layer of the sound data processing model, to obtain first acoustic characteristic information corresponding to each piece of sound data, the first acoustic characteristic information being an acoustic characteristic with a time scale less than or equal to a preset time scale.

[0055] The first acoustic characteristic information can represent short-time or instantaneous acoustic characteristics, and can include at least one of pitch, sound energy, mel-frequency cepstral coefficient, linear predictive coding, cepstral coefficient of linear predictive coding, fundamental frequency, and formant; the sound energy includes at least one of original energy, energy dynamic range, and energy envelope shape.

[0056] Preferably, the first acoustic characteristic information can also include the instantaneous change rate (first derivative) of the above-mentioned acoustic characteristics and the change rate (second derivative) of the instantaneous change rate, to better represent short-time acoustic characteristics.

[0057] S230, the second acoustic feature extraction layer based on the sound data processing model extracts sound features from the candidate sound information, and obtains second acoustic feature information corresponding to each piece of sound data; the candidate sound information is each piece of sound data or first acoustic feature information corresponding to each piece of sound data; the second acoustic feature information is an acoustic feature with a time scale greater than a preset time scale; the first acoustic feature information and the second acoustic feature information are mutually decoupled from the content in the audio data set.

[0058] The second acoustic feature information can represent long-time context features at different time scales. The second acoustic feature information can include a distribution of third acoustic feature information and / or a change trend of the third acoustic feature information; the third acoustic feature information is a plurality of first acoustic feature information within a time range corresponding to the second acoustic feature information. The first acoustic feature information and the second acoustic feature information are mutually decoupled from the content in the audio data set, which can effectively avoid the influence of the text content of the sound data on the timbre itself.

[0059] In the embodiments of the present application, the first acoustic feature extraction layer and the second acoustic feature extraction layer are connected in parallel or in series.

[0060] When the first acoustic feature extraction layer and the second acoustic feature extraction layer are connected in parallel, the candidate sound information is each piece of sound data, and the first acoustic feature information and the second acoustic feature information can be obtained simultaneously and side by side, that is, the first acoustic feature extraction layer and the second acoustic feature extraction layer of the sound data processing model simultaneously extract sound features from each piece of sound data in the audio data set to obtain the first acoustic feature information and the second acoustic feature information corresponding to each piece of sound data.

[0061] When the first acoustic feature extraction layer and the second acoustic feature extraction layer are connected in series, the candidate sound information is the first acoustic feature information corresponding to each piece of sound data, and the first acoustic feature information and the second acoustic feature information are obtained in the following process: first, the first acoustic feature extraction layer of the sound data processing model extracts sound features from each piece of sound data in the audio data set to obtain the first acoustic feature information corresponding to each piece of sound data; then, the first acoustic feature information is input into the second acoustic feature extraction layer of the sound data processing model, and the second acoustic feature extraction layer of the sound data processing model analyzes the first acoustic feature information corresponding to each piece of sound data to obtain the second acoustic feature information corresponding to each piece of sound data; the feature data analysis includes data distribution and / or data distribution trend.

[0062] The data distribution situation can describe the value range, central tendency, dispersion degree or morphological characteristics of the data on the static dimension, and reflect the overall state of the data. The static dimension can represent a dimension that does not involve time change or sequence change.

[0063] The common indicators or morphologies corresponding to the data distribution situation can include: data central tendency: mean, median, mode. Data dispersion degree: variance, standard deviation, interquartile range. Data distribution morphology: normal distribution, skew distribution, uniform distribution, bimodal distribution, etc. The data distribution morphology can be analyzed and obtained through visualization means such as histogram and box plot.

[0064] The data distribution trend can be understood as describing the increase-decrease law, fluctuation characteristics or development direction of the data on the dynamic dimension, and reflects the process of the data "changing with a certain variable". The dynamic dimension can be the dimension of the change law of the data with time or sequence, emphasizing dynamic change. The dynamic dimension can usually include time dimension and ordered dimension.

[0065] The common indicators or morphologies corresponding to the data distribution trend can include: data trend direction: upward, downward, stable. Data fluctuation characteristics: periodic fluctuation (such as change in each unit of time), random fluctuation, linear or nonlinear growth. Data change rate: growth rate, slope.

[0066] For example, the second acoustic feature information can be regarded as a statistical value extracted and analyzed from the first acoustic feature information in the current window through different lengths of sliding windows. The length of the sliding window is aligned with each time scale, for example, the window W1 length is 10ms, the window W2 length is 20ms, etc. For each kind of first acoustic feature information under each window length, the numerical value of the first acoustic feature information is obtained, and the statistical quantity such as mean value is calculated, for example, the window W1 length is 10ms, and its coverage range is 0ms-10ms in the second acoustic feature information capture. Assuming that each first acoustic feature information is 1ms, then the window W1 will calculate the statistical quantity of the pitch and the statistical quantity of the energy at 1, 2, …, 10ms, respectively, to form the second acoustic feature information under the window W1. If the time scale is 20ms, the window W2 length is 20ms, and the statistical quantity of various first acoustic feature information in 0-20s is calculated.

[0067] S240, processing the first acoustic feature information and the second acoustic feature information based on the sound data processing model to obtain at least one sound clustering information.

[0068] Specifically, the first acoustic feature information and the second acoustic feature information of each piece of sound data are mapped into the latent space as a latent vector of the first acoustic feature information and a latent vector of the second acoustic feature information, then the latent vector of the first acoustic feature information and the latent vector of the second acoustic feature information are fused in the latent space to obtain a latent space fusion vector, and then clustering processing is performed based on the latent space fusion vector to obtain sound clustering information. The fusion processing mode can include one of latent vector splicing, latent vector averaging, latent vector summation, and latent vector weighted summation.

[0069] The technical scheme of the embodiment determines an audio data set and a sound data processing model; the audio data set contains at least one piece of sound data; the sound data processing model includes an acoustic feature extraction layer, and the acoustic feature extraction layer includes a first acoustic feature extraction layer and a second acoustic feature extraction layer. The first acoustic feature extraction layer of the sound data processing model is used to perform sound feature extraction on each piece of sound data in the audio data set, so as to accurately obtain the first acoustic feature information corresponding to each piece of sound data. Further, the second acoustic feature extraction layer of the sound data processing model is used to perform sound feature extraction on candidate sound information, so as to obtain the second acoustic feature information corresponding to each piece of sound data; the candidate sound information is each piece of sound data or the first acoustic feature information corresponding to each piece of sound data, so that the second acoustic feature extraction layer of the sound data processing model can perform sound feature extraction on different data, thereby obtaining accurate second acoustic feature information in different time scales, so as to facilitate subsequent processing of the first acoustic feature information and the second acoustic feature information based on the sound data processing model, and improve the accuracy of at least one piece of sound clustering information obtained. In addition, the first acoustic feature information and the second acoustic feature information are decoupled from the content in the audio data set, so that the influence of the text content in the sound data on the timbre itself can be effectively avoided.

[0070] Embodiment three

[0071] Figure 3 A flowchart of another sound data processing method provided by the embodiment of the application is provided. The technical scheme of the embodiment is further optimized based on the process of S120 in the foregoing embodiment, and the embodiment can be combined with each optional scheme in one or more of the foregoing embodiments. As shown in Figure 3 The sound data processing method provided by the application includes the following processes:

[0072] S310, determining an audio data set and a sound data processing model; the audio data set contains at least one piece of sound data; the sound data processing model includes a first encoder, at least one second encoder, an acoustic feature extraction layer, a fusion layer, and a clustering layer, different dimensions of second acoustic feature information match different second encoders, and the acoustic feature extraction layer includes a first acoustic feature extraction layer and a second acoustic feature extraction layer.

[0073] The acoustic feature extraction layer is configured to perform acoustic feature extraction on each piece of sound data in the audio data set; the fusion layer is configured to map the acoustic features to a latent space by using an acoustic feature matching encoder to obtain latent vectors of the acoustic features, and perform feature fusion on the latent vectors of the acoustic features in the latent space; and the clustering layer is configured to perform clustering processing on the data after the feature fusion to obtain sound clustering information.

[0074] The second acoustic feature information of different dimensions matches different second encoders, and each second encoder has the same structure, but the parameters of the second encoders are different due to different time scales of the features represented by the vectors. For example, the second acoustic feature information with a shorter time scale focuses more on the microscopic dynamic changes of the sound data in a short period of time, and can reflect relatively more local acoustic features; however, the second acoustic feature information with a longer time scale focuses more on the macroscopic trend of the sound, and can reflect more stable acoustic features. However, it is difficult for one encoder to capture both instantaneous features and macroscopic trends, so different encoders are used for different time scales, so that the feature extraction capability is stronger and more stable.

[0075] S320, obtaining the first acoustic feature information and the second acoustic feature information corresponding to each piece of sound data, controlling the first encoder to map the first acoustic feature information to a latent space to obtain at least one first latent vector, and controlling the second encoder to map the second acoustic feature information to the latent space to obtain at least one second latent vector.

[0076] Specifically, the first latent vector and the second latent vector corresponding to each piece of sound data are mapped into the same latent space.

[0077] S330, based on the fusion layer, performing feature fusion on the first latent vector and the second latent vector in the latent space by using a preset fusion method to obtain feature fusion data of each piece of sound data.

[0078] Specifically, the first latent vector and the second latent vector are spliced in the same latent space to obtain the feature fusion data of each piece of sound data.

[0079] Optionally, the preset fusion method can include one of a first fusion method, a second fusion method, and a third fusion method; the first fusion method is a method of averaging acoustic feature information; the second fusion method is a method of summing acoustic feature information; and the third fusion method is a method of weighted summing acoustic feature information.

[0080] The preset fusion method is a first fusion method. The first hidden vector and the second hidden vector are fused in the hidden space based on the fusion layer by using the preset fusion method to obtain feature fusion data of each piece of sound data, including: summing all the first hidden vectors and all the second hidden vectors to obtain a reference hidden vector, and dividing the reference hidden vector by the total number of the first hidden vectors and the second hidden vectors corresponding to each piece of sound data to obtain the feature fusion data of each piece of sound data.

[0081] The preset fusion method can be a second fusion method. The first hidden vector and the second hidden vector are fused in the hidden space based on the fusion layer by using the preset fusion method to obtain feature fusion data of each piece of sound data, which can include: summing all the first hidden vectors and all the second hidden vectors corresponding to each piece of sound data to obtain the feature fusion data of each piece of sound data.

[0082] The preset fusion method is preferably a third fusion method. The fusion layer of the sound data processing model includes a full connection layer. The first hidden vector and the second hidden vector are fused in the hidden space based on the fusion layer by using the preset fusion method to obtain feature fusion data of each piece of sound data, which can include the following steps B1-B2:

[0083] Step B1, based on the full connection layer, the element values of each dimension of the first hidden vector and the element values of each dimension of the second hidden vector, determining the first weight of the first hidden vector and the second weight of the second hidden vector.

[0084] Specifically, the first weight of the first hidden vector is determined based on the full connection layer and the element values of each dimension of the first hidden vector, and the first weight of the second hidden vector is determined based on the full connection layer and the element values of each dimension of the second hidden vector.

[0085] In this embodiment, optionally, based on the full connection layer, the element values of each dimension of the first hidden vector and the element values of each dimension of the second hidden vector, determining the first weight of the first hidden vector and the second weight of the second hidden vector can include: determining the first weight of the first hidden vector based on the full connection layer, the element values of each dimension of the first hidden vector and a first preset bias term; determining the second weight of the second hidden vector based on the full connection layer, the element values of each dimension of the second hidden vector and a second preset bias term; and normalizing all the first weights of the first hidden vectors and all the second weights of the second hidden vectors corresponding to the same piece of sound data to obtain the updated first weight of the first hidden vector and the updated second weight of the second hidden vector.

[0086] Specifically, based on the element values of each dimension of the first hidden vector and the full connection layer, a first reference weight corresponding to the element values of each dimension of the first hidden vector is mapped, the element values of each dimension of the first hidden vector and the first reference weight corresponding to the element values of each dimension of the first hidden vector are multiplied to obtain a second reference weight, and the first preset bias is added to the second reference weight to obtain the first weight of the first hidden vector.

[0087] For example, the first hidden vector [h1, h2, …, hN], the first reference weight [w1, w2, …, wN], the first preset bias b, and the first weight R of the first hidden vector are obtained by w1×h1+w2×h2+…+wN×hN+b. n n n n

[0088] Similarly, based on the element values of each dimension of the second hidden vector and the full connection layer, a third reference weight corresponding to the element values of each dimension of the second hidden vector is mapped, the element values of each dimension of the second hidden vector and the third reference weight corresponding to the element values of each dimension of the second hidden vector are multiplied to obtain a fourth reference weight, and the second preset bias is added to the fourth reference weight to obtain the second weight of the second hidden vector.

[0089] The technical scheme of the embodiment of the application introduces the first preset bias and the second preset bias to realize more accurate determination of the weight of the hidden vector and avoid low accuracy of the determined weight due to errors. Further, the first weight of all first hidden vectors and the second weight of all second hidden vectors corresponding to the same sound data are normalized to obtain the first weight of the updated first hidden vector and the second weight of the updated second hidden vector. The normalization processing ensures that the sum of the weights of all hidden vectors of the same sound data is one, so that subsequent weighted fusion of the first hidden vector and the second hidden vector based on the first hidden vector, the second hidden vector, the first weight and the second weight can obtain more accurate feature fusion data of the same sound data.

[0090] Step B2, based on the first hidden vector, the second hidden vector, the first weight and the second weight, the first hidden vector and the second hidden vector are weighted and fused to determine the feature fusion data of the same sound data.

[0091] Specifically, the first hidden vector and the first weight of the first hidden vector are multiplied to obtain a first fusion vector, the second hidden vector and the second weight of the second hidden vector are multiplied to obtain a second fusion vector, all first fusion vectors and second fusion vectors are summed to obtain a third fusion vector, and then the third fusion vector is divided by the total number of the first hidden vector and the second hidden vector to obtain the feature fusion data of the same sound data. ​​​​

[0092] The full connection layer of the sound data processing model ensures rapid and accurate determination of the weight of each latent vector, and the fusion layer of the sound data processing model further adopts a weighted fusion manner to realize accurate determination of the quantized feature fusion data of the same sound data.

[0093] In S340, the preset clustering method is used to cluster all the feature fusion data to obtain at least one sound clustering information.

[0094] The preset clustering method can be a clustering method based on spatial distance, and the calculation method of the spatial distance can include cosine similarity, Euclidean distance, etc.

[0095] Specifically, the preset clustering method is used to determine the reference spatial distance of each sound data based on the comparison result of the reference spatial distance and the maximum clustering distance threshold, and the clustering result of each sound data is determined, specifically: when the reference spatial distance is less than or equal to the maximum clustering distance threshold, the corresponding acoustic feature information of the feature fusion data belongs to the same object; when the reference spatial distance is greater than the maximum clustering distance threshold, the corresponding acoustic feature information of the feature fusion data does not belong to the same object.

[0096] The technical scheme of the embodiment of the application determines an audio data set and a sound data processing model; the audio data set contains at least one sound data; the sound data processing model includes a first encoder, at least one second encoder, an acoustic feature extraction layer, a fusion layer, and a clustering layer, different dimensions of second acoustic feature information match different second encoders, the acoustic feature extraction layer includes a first acoustic feature extraction layer and a second acoustic feature extraction layer. The first acoustic feature information and the second acoustic feature information corresponding to each sound data are obtained, the first encoder is controlled to map the first acoustic feature information to a latent space to obtain at least one first latent vector, and the second encoder is controlled to map the second acoustic feature information to the latent space to obtain at least one second latent vector; the dimension of the latent vector is lower, which can ensure that the subsequent calculation process is faster, i.e., the calculation efficiency is improved. Further, based on the fusion layer, the first latent vector and the second latent vector are fused in the latent space by using a preset fusion method to obtain feature fusion data of each sound data, the fusion effect in the latent space is more natural and stable than that in the explicit space, so that the preset clustering method is used to cluster all the feature fusion data based on the clustering layer to obtain at least one sound clustering information, which is more accurate.

[0097] Embodiment four

[0098] Figure 4A structural schematic diagram of a sound data processing apparatus provided by an embodiment of the present application can be applicable to the case of processing sound data. The sound data processing apparatus can be realized in the form of hardware and / or software, and can be configured in any electronic device with network communication function. As shown in Figure 4 the sound data processing apparatus of the present application comprises:

[0099] a data set determination module 410 configured to determine an audio data set containing at least one piece of sound data;

[0100] a processing module 420 configured to process the audio data set based on a sound data processing model to obtain at least one piece of sound clustering information; each piece of sound clustering information contains acoustic feature information of the same object; wherein the sound data processing model is used to convert each piece of sound data into an embedding space fusion vector, and perform clustering processing based on the embedding space fusion vector to obtain sound clustering information.

[0101] On the basis of the above-mentioned embodiment, the processing module can include a first data processing unit, which is configured to: perform sound feature extraction on each piece of sound data in the audio data set based on a first acoustic feature extraction layer of the sound data processing model to obtain first acoustic feature information corresponding to each piece of sound data, the first acoustic feature information being acoustic features with a time scale less than or equal to a preset time scale; perform sound feature extraction on candidate sound information based on a second acoustic feature extraction layer of the sound data processing model to obtain second acoustic feature information corresponding to each piece of sound data; the candidate sound information being each piece of sound data or the first acoustic feature information corresponding to each piece of sound data; the second acoustic feature information being acoustic features with a time scale greater than the preset time scale; and the first acoustic feature information and the second acoustic feature information being mutually decoupled from the content in the audio data set.

[0102] On the basis of the above-mentioned embodiment, the sound data processing model can include a first encoder, at least one second encoder, a fusion layer and a clustering layer, and different dimensions of the second acoustic feature information match different second encoders.

[0103] The processing module comprises a second data processing unit, which is configured to: control the first encoder to map the first acoustic feature information to a latent space to obtain at least one first latent vector, and control the second encoder to map the second acoustic feature information to the latent space to obtain at least one second latent vector; based on the fusion layer, perform feature fusion on the first latent vector and the second latent vector in the latent space by using a preset fusion method to obtain feature fusion data of each piece of sound data; and based on the clustering layer, perform clustering processing on all the feature fusion data by using a preset clustering method to obtain at least one sound clustering information.

[0104] On the basis of the above-mentioned embodiments, the first data processing unit is further configured to: based on the second acoustic feature extraction layer of the sound data processing model, perform feature data analysis on the first acoustic feature information corresponding to each piece of sound data to obtain second acoustic feature information corresponding to each piece of sound data; and the feature data analysis comprises data distribution and / or data distribution trend.

[0105] On the basis of the above-mentioned embodiments, the preset fusion method comprises one of a first fusion method, a second fusion method and a third fusion method; the first fusion method is a method of averaging acoustic feature information; the second fusion method is a method of summing acoustic feature information; and the third fusion method is a method of weighted summing acoustic feature information.

[0106] On the basis of the above-mentioned embodiments, the preset fusion method is the third fusion method; the fusion layer of the sound data processing model comprises a full connection layer; and the processing module comprises a third data processing unit, which is configured to: based on the full connection layer, element values of each dimension of the first latent vector and element values of each dimension of the second latent vector, determine a first weight of the first latent vector and a second weight of the second latent vector; and based on the first latent vector, the second latent vector, the first weight and the second weight, perform weighted fusion on the first latent vector and the second latent vector to determine the feature fusion data of the same piece of sound data.

[0107] On the basis of the above-mentioned embodiments, the third data processing unit is further configured to: based on the full connection layer, element values of each dimension of the first latent vector and a first preset bias term, determine the first weight of the first latent vector; based on the full connection layer, element values of each dimension of the second latent vector and a second preset bias term, determine the second weight of the second latent vector; and perform normalization processing on the first weight of all the first latent vectors and the second weight of all the second latent vectors corresponding to the same piece of sound data to obtain an updated first weight of the first latent vector and an updated second weight of the second latent vector.

[0108] In the above embodiment, optionally, the first acoustic feature information comprises at least one of pitch, sound energy, mel-frequency cepstral coefficient, linear predictive coding, cepstral coefficient of linear predictive coding, fundamental frequency and formant; and the sound energy comprises at least one of raw energy, energy dynamic range and energy envelope shape.

[0109] The second acoustic feature information comprises distribution of third acoustic feature information and / or variation trend of the third acoustic feature information; and the third acoustic feature information is the first acoustic feature information in a time range corresponding to the second acoustic feature information.

[0110] In the above embodiment, optionally, the processing module comprises a fourth data processing unit, configured to: determine the latent space distance between the latent space fusion vectors of the sound data; and determine the clustering result of the sound data based on a comparison result of the latent space distance and a maximum clustering distance threshold.

[0111] The sound data processing apparatus provided in the embodiments of the present application can execute the sound data processing method provided in any of the embodiments of the present application, and has the corresponding function modules and beneficial effects of the execution method.

[0112] Embodiment five

[0113] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0114] Figure 5 A structural schematic diagram of an electronic device that can be used to implement the sound data processing method of the embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present application described and / or claimed in this document.

[0115] As Figure 5As shown, the electronic device 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., communicatively connected to the at least one processor 11, where the memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0116] Various components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc., an output unit 17, such as various types of displays, a speaker, etc., a storage unit 18, such as a magnetic disk, an optical disk, etc., and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0117] The processor 11 can be various general and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs various methods and processes described above, such as the sound data processing method.

[0118] In some embodiments, the sound data processing method can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the sound data processing method described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the sound data processing method by any other appropriate means, such as by means of firmware.

[0119] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a special-purpose standard product (ASSP), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0120] Computer programs used to implement the methods of the application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program, when executed, implements the functions / acts specified in the flowcharts and / or block diagrams. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as a standalone software package, and partially on a remote machine or entirely on a remote machine or server.

[0121] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store computer programs for use by or in connection with an instruction execution system, apparatus, or device. Computer-readable storage media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0122] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0123] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0124] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.

[0125] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be executed in parallel, executed in sequence, or executed in a different order, as long as the desired results of the present disclosure are achieved, and the present disclosure is not limited herein.

[0126] The specific embodiments described above are not intended to be limiting, and persons skilled in the art will appreciate that various modifications, combinations, sub-combinations and alternatives can be made to the specific embodiments without departing from the spirit and principles of the disclosure. Accordingly, the disclosure is not limited to the specific embodiments described above, but only by the scope of the appended claims.

Claims

1. A sound data processing method characterized by comprising: The method comprises: determining an audio data set containing at least one piece of sound data; processing the audio data set based on a sound data processing model to obtain at least one piece of sound clustering information; each piece of sound clustering information contains acoustic feature information of the same object; wherein the sound data processing model is used to convert each piece of sound data into an implicit space fusion vector, and perform clustering processing based on the implicit space fusion vector to obtain sound clustering information; the processing of the audio data set based on the sound data processing model comprises: performing sound feature extraction on each piece of sound data in the audio data set based on a first acoustic feature extraction layer of the sound data processing model to obtain first acoustic feature information corresponding to each piece of sound data, wherein the first acoustic feature information is acoustic features with a time scale less than or equal to a preset time scale; performing sound feature extraction on candidate sound information under different lengths of time windows based on a second acoustic feature extraction layer of the sound data processing model to obtain at least one piece of second acoustic feature information corresponding to each piece of sound data; the candidate sound information is each piece of sound data or the first acoustic feature information corresponding to each piece of sound data; the second acoustic feature information is acoustic features with a time scale greater than the preset time scale, used to represent long-time context features under different time scales; the first acoustic feature information and the second acoustic feature information are mutually decoupled from the content in the audio data set.

2. The method of claim 1, wherein, The sound data processing model comprises a first encoder, at least one second encoder, a fusion layer, and a clustering layer, and different dimensions of the second acoustic feature information match different second encoders; correspondingly, after obtaining the first acoustic feature information and the second acoustic feature information corresponding to each piece of sound data, the method further comprises: controlling the first encoder to map the first acoustic feature information to an implicit space to obtain at least one first implicit vector, and controlling the second encoder to map the second acoustic feature information to the implicit space to obtain at least one second implicit vector; based on the fusion layer, performing feature fusion on the first implicit vector and the second implicit vector in the implicit space using a preset fusion method to obtain feature fusion data of each piece of sound data; based on the clustering layer, performing clustering processing on all the feature fusion data using a preset clustering method to obtain at least one piece of sound clustering information.

3. The method of claim 1, wherein, performing sound feature extraction on candidate sound information based on the second acoustic feature extraction layer of the sound data processing model to obtain second acoustic feature information corresponding to each piece of sound data, comprising: performing feature data analysis on the first acoustic feature information corresponding to each piece of sound data based on the second acoustic feature extraction layer of the sound data processing model to obtain second acoustic feature information corresponding to each piece of sound data; the feature data analysis includes data distribution and / or data distribution trend.

4. The method of claim 2, wherein, The preset fusion method includes one of a first fusion method, a second fusion method, and a third fusion method; the first fusion method is a method of averaging acoustic feature information; the second fusion method is a method of summing acoustic feature information; and the third fusion method is a method of weighted summing acoustic feature information.

5. The method of claim 4, wherein, The preset fusion method is the third fusion method; and the fusion layer of the sound data processing model includes a full connection layer. Correspondingly, based on the fusion layer, the first hidden vector and the second hidden vector are fused in the hidden space by using the preset fusion method to obtain feature fusion data of each piece of sound data, including: Based on the full connection layer, element values of each dimension of the first hidden vector, and element values of each dimension of the second hidden vector, a first weight of the first hidden vector and a second weight of the second hidden vector are determined; Based on the first hidden vector, the second hidden vector, the first weight, and the second weight, the first hidden vector and the second hidden vector are weighted fused to determine the feature fusion data of the same piece of sound data.

6. The method of claim 5, wherein, Based on the full connection layer, element values of each dimension of the first hidden vector, and element values of each dimension of the second hidden vector, a first weight of the first hidden vector and a second weight of the second hidden vector are determined, including: Based on the full connection layer, element values of each dimension of the first hidden vector, and a first preset bias term, a first weight of the first hidden vector is determined; Based on the full connection layer, element values of each dimension of the second hidden vector, and a second preset bias term, a second weight of the second hidden vector is determined; The first weight of the first hidden vector and the second weight of the second hidden vector corresponding to the same piece of sound data are normalized to obtain an updated first weight of the first hidden vector and an updated second weight of the second hidden vector.

7. The method of claim 1, wherein, The first acoustic feature information includes at least one of pitch, sound energy, mel-frequency cepstral coefficient, linear predictive coding, cepstral coefficient of linear predictive coding, fundamental frequency, and formant; and the sound energy includes at least one of original energy, energy dynamic range, and energy envelope shape. The second acoustic feature information includes a distribution of third acoustic feature information and / or a change trend of the third acoustic feature information; and the third acoustic feature information is first acoustic feature information in a time range corresponding to the second acoustic feature information.

8. The method of claim 1, wherein, Based on the hidden space fusion vector, clustering processing is performed to obtain sound clustering information, including: A hidden space distance between the hidden space fusion vectors of each piece of sound data is determined; Based on a comparison result of the hidden space distance and a maximum clustering distance threshold, a clustering result of each piece of sound data is determined.

9. A sound data processing apparatus characterized by comprising: The device includes: A data set determination module is configured to determine an audio data set, and the audio data set includes at least one piece of sound data; The processing module is configured to process the audio data set based on a sound data processing model to obtain at least one sound cluster information; each sound cluster information contains acoustic feature information of a same object; wherein the sound data processing model is configured to convert each piece of sound data into an implicit space fusion vector, and perform clustering processing based on the implicit space fusion vector to obtain the sound cluster information; The processing module includes a first data processing unit, which is configured to: perform sound feature extraction on each piece of sound data in the audio data set based on a first acoustic feature extraction layer of a sound data processing model to obtain first acoustic feature information corresponding to each piece of sound data, the first acoustic feature information being acoustic features with a time scale less than or equal to a preset time scale; perform sound feature extraction on candidate sound information under different lengths of time windows based on a second acoustic feature extraction layer of the sound data processing model to obtain at least one second acoustic feature information corresponding to each piece of sound data; the candidate sound information being each piece of sound data or the first acoustic feature information corresponding to each piece of sound data; the second acoustic feature information being acoustic features with a time scale greater than the preset time scale, used to represent long-time context features under different time scales; and the first acoustic feature information and the second acoustic feature information being mutually decoupled from the content in the audio data set.

Citation Information

Patent Citations

  • Method and device for identifying authenticity of sound in audio information

    CN120148554A