A data processing method, apparatus, device, storage medium, and program product
By adjusting the parameters of the initial accent classification model, using the differences in accent characteristics of multi-snippet speech information and the differences in the types of accents to be determined and the types of accents in the sample, the problem of low accent classification accuracy in the prior art is solved, and high-precision speech recognition is achieved.
Patent Information
- Application Number
- CN202111173239.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-08
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-10-08
AI Technical Summary
In the prior art, the classification of voice information based on the accent characteristics is relatively low, and it is difficult to meet the high-precision voice recognition needs.
By obtaining the speech training samples of the corresponding accent type of the sample, determining the accent features and the to-determined accent types corresponding to the multi-snippet speech information, defining the first loss function and the second loss function, so as to adjust the parameters of the initial accent classification model and improving the accuracy and rationality of accent classification.
The trained model can output more accurate accent classification results and can reasonably extract accent features of speech information, improving the accuracy and rationality of the model.
Smart Images

Figure CN114328811B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data models, and in particular, to a data processing method, apparatus, device, storage medium, and program product. Background Art
[0002] In real life, everyone has their own unique accent when speaking. For example, people from different countries may have the accent characteristics of their own languages when speaking Chinese, thus speaking Chinese with different accents.
[0003] In order to achieve more accurate speech recognition, in the related art, the collected speech information can be classified based on accent characteristics first, and the speech information with different accent characteristics is input into different speech recognition models for speech content recognition. It can be seen that the accuracy of accent classification will directly affect the accuracy of speech recognition.
[0004] However, the accuracy of classifying speech information based on accent characteristics in the related art is low, and it is difficult to meet the requirements of high-precision speech recognition. Summary of the Invention
[0005] To solve the above technical problems, this application provides a data processing method, apparatus, device, storage medium, and program product, so that the trained model can not only output a relatively accurate accent classification result, but also extract accent features from speech information more reasonably, improving the rationality and accuracy of the model.
[0006] The embodiments of this application disclose the following technical solutions:
[0007] In a first aspect, the embodiments of this application disclose a data processing method, and the method includes:
[0008] Obtain a first speech training sample corresponding to a sample accent type, where the first speech training sample includes multiple sub-speech information of users corresponding to the same accent type;
[0009] Determine the accent features corresponding to the multiple sub-speech information respectively according to the first speech training sample and an initial accent classification model, and a to-be-determined accent type corresponding to the first speech training sample, where the to-be-determined accent type is determined based on the accent features;
[0010] Determine a first loss function and a second loss function corresponding to the first speech training sample, where the first loss function is determined according to the differences between the accent features corresponding to the multiple sub-speech information respectively, and the second loss function is determined according to the difference between the to-be-determined accent type and the sample accent type;
[0011] Adjust the parameters of the initial accent classification model according to the first loss function and the second loss function to obtain an accent classification model, which is used to determine the accent type corresponding to the speech information to be classified.
[0012] In a second aspect, an embodiment of the present application discloses a data processing device, which includes a first acquisition unit, a first determination unit, a second determination unit, and a first adjustment unit:
[0013] The first acquisition unit is configured to acquire a first speech training sample corresponding to a sample accent type, and the first speech training sample includes multiple sub-speech information of users with the same accent type;
[0014] The first determination unit is configured to determine the accent features corresponding to the multiple sub-speech information respectively according to the first speech training sample and the initial accent classification model, and the pending accent type corresponding to the first speech training sample, where the pending accent type is determined based on the accent features;
[0015] The second determination unit is configured to determine a first loss function and a second loss function corresponding to the first speech training sample, where the first loss function is determined according to the differences between the accent features corresponding to the multiple sub-speech information respectively, and the second loss function is determined according to the difference between the pending accent type and the sample accent type;
[0016] The first adjustment unit is configured to adjust the parameters of the initial accent classification model according to the first loss function and the second loss function to obtain an accent classification model, which is used to determine the accent type corresponding to the speech information to be classified.
[0017] In a possible implementation manner, the first speech training sample is the same as the second speech training sample.
[0018] In a possible implementation manner, the first speech training sample is any one of multiple first speech training samples included in a speech training sample set, and the multiple sub-speech information included in the multiple first speech training samples corresponds to a first user and a second user, and the accent type of the first user is different from that of the second user.
[0019] In a possible implementation manner, the multiple first speech training samples correspond to the same language, and users with different accent types correspond to different accent regions.
[0020] In a third aspect, an embodiment of the present application discloses a computer device, which includes a processor and a memory:
[0021] The memory is used to store program code and transmit the program code to the processor;
[0022] The processor is used to execute the data processing method described in the first aspect according to the instructions in the program code.
[0023] In a fourth aspect, an embodiment of the present application discloses a computer-readable storage medium, which is used to store a computer program, and the computer program is used to execute the data processing method described in the first aspect.
[0024] In a fifth aspect, an embodiment of the present application discloses a computer program product including instructions, which, when running on a computer, causes the computer to execute the data processing method described in any item of the first aspect
[0025] It can be seen from the above technical solutions that when training the initial accent classification model, since the multiple sub-speech information included in the first speech training sample corresponds to users of the same accent type, the multiple sub-speech information should have similar accent characteristics under normal circumstances. Based on this, after determining the accent characteristics corresponding to the multiple sub-speech information and the pending accent type corresponding to the first sample speech information through the initial accent classification model, the first loss function and the second loss function corresponding to the first speech training sample can be determined. The first loss function is determined according to the differences between the accent characteristics corresponding to the multiple sub-speech information, and the second loss function is determined according to the difference between the pending accent type and the sample accent type. Therefore, by adjusting the parameters of the initial accent classification model based on the first loss function and the second loss function, on the one hand, the accent type determined by the initial accent classification model can be made more accurate, and on the other hand, during the training process of the model, the differences between the accent characteristics of the determined sub-speech information can be controlled within a reasonable range, making the way of determining the accent characteristics more in line with the real accent situation, and improving the accuracy and rationality of the training of the accent classification model. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0027] Figure 1 A schematic diagram of an accent classification scenario provided by an embodiment of the present application;
[0028] Figure 2Schematic diagram of a data processing method in an actual application scenario provided by an embodiment of the present application;
[0029] Figure 3 Flowchart of a data processing method provided by an embodiment of the present application;
[0030] Figure 4 Schematic diagram of a model training provided by an embodiment of the present application;
[0031] Figure 5 Schematic diagram of an initial accent classification model provided by an embodiment of the present application;
[0032] Figure 6 Flowchart of a data processing method in an actual application scenario provided by an embodiment of the present application;
[0033] Figure 7 Schematic diagram of a model provided by an embodiment of the present application;
[0034] Figure 8 Schematic diagram of an application scenario provided by an embodiment of the present application;
[0035] Figure 9 Structural block diagram of a data processing device provided by an embodiment of the present application;
[0036] Figure 10 Structural diagram of a computer device provided by an embodiment of the present application;
[0037] Figure 11 Structural diagram of a server provided by an embodiment of the present application. Detailed implementation manners
[0038] The embodiments of the present application will be described below with reference to the accompanying drawings.
[0039] Speech recognition technology involves all aspects of people's lives. For example, through speech recognition, people can control smart homes to work by speaking, control the voice interaction device on the vehicle to play music, and broadcast weather and traffic information. Among them, the prerequisite for the use of the voice control method is the ability to accurately recognize the content of the voice. In real life, different users often have different accents when speaking. For example, even when speaking Chinese, people from different provinces or different countries may have different accents when speaking Chinese. Therefore, in order to perform more accurate speech recognition, the speech information can be classified based on the accent characteristics first, and the speech information with different accent characteristics can be recognized in a targeted manner. For example Figure 1 As shown, after obtaining the speech to be recognized, it can be first input into the accent classifier to determine the accent type corresponding to the speech to be recognized, and then input into the corresponding model according to the accent type for speech recognition.
[0040] In the related art, when performing accent classification, usually, each segment of the speech to be recognized is first analyzed by an accent classification model to obtain the corresponding accent features, and then the accent features corresponding to multiple segments of speech are averaged, and the processed result is used as the overall accent type of the speech to be recognized. However, in the process of training the accent classification model, due to only focusing on the accuracy of the finally determined accent type and ignoring the rationality of determining the accent features of each segment of speech, it may lead to the overfitting problem of the accent classification model with respect to the training sample set, that is, although it has a high classification accuracy for the training sample set, for multiple segments of speech of users with the same accent type, it may determine quite different accent features, which does not conform to the characteristic that the speeches of users with the same accent type usually have similar accent features in the real situation. As a result, the model cannot accurately and reasonably extract accent features and is difficult to be applied to actual accent classification.
[0041] To solve the above technical problems, an embodiment of the present application provides a data processing method. When a processing device trains an accent classification model, on the one hand, multiple loss functions can be used to constrain the difference between the to-be-determined accent type determined by the model and the sample accent type; on the other hand, the difference between the accent features determined by the model for multiple sub-speech information of users with the same accent type can be constrained, so as to conform to the accent characteristics of users with the same accent type in the real situation, enabling the trained model to not only output a relatively accurate accent classification result but also extract accent features from speech information more reasonably, improving the rationality and accuracy of the model.
[0042] It can be understood that this method can be applied to a processing device, which is a processing device capable of data processing, such as a terminal device or a server with data processing functions. This method can be independently executed by the terminal device or the server, or can be applied to a network scenario where the terminal device and the server communicate and is executed in cooperation by the terminal device and the server. Among them, the terminal device can be a device such as a computer or a mobile phone. The server can be understood as an application server or a Web server. In actual deployment, the server can be an independent server or a cluster server.
[0043] In addition, this application also relates to Artificial Intelligence (AI) technology. Artificial Intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, Artificial Intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial Intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.
[0044] Artificial Intelligence technology is an interdisciplinary subject that involves a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of Artificial Intelligence generally include technologies such as sensors, dedicated Artificial Intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of Artificial Intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning. This application mainly relates to speech technology, natural language processing technology, and machine learning technology among them.
[0045] The key technologies of Speech Technology include Automatic Speech Recognition (ASR) technology, Text-to-Speech (TTS) technology, and voiceprint recognition technology. Enabling computers to listen, see, speak, and feel is the future development direction of human-computer interaction, and among them, speech has become one of the most promising human-computer interaction methods in the future.
[0046] Natural Language Processing (NLP) is an important direction in the field of computer science and Artificial Intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural Language Processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language used by people in daily life, so it has a close connection with the research of linguistics. Natural Language Processing technology usually includes technologies such as text processing, semantic understanding, machine translation, robot question answering, and knowledge graphs.
[0047] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0048] In the embodiments of the present application, the extraction of accent features requires the use of speech technology and natural language processing technology; in the process of model training, the loss function is used to adjust the model parameters, and machine learning technology can be used in the pre-training process of some models.
[0049] To facilitate the understanding of the technical solutions provided by the embodiments of the present application, next, a data processing method provided by the embodiments of the present application will be introduced in combination with an actual application scenario.
[0050] See Figure 2 , Figure 2 is a schematic diagram of a data processing method in an actual application scenario provided by the embodiments of the present application. In this actual application scenario, the processing device is server 201.
[0051] In server 201, there is an initial accent classification model, which can be used to classify the accent types of speech information. When training the initial accent classification model, server 201 can first obtain the first speech training samples corresponding to the sample accent types. The first speech training samples include multiple sub-speech information corresponding to users A, B, and C. For example, the multiple sub-speech information can be the speech information corresponding to a section of speech of the three users at different time periods. Among them, these three users have the same accent type, that is, they all correspond to the sample accent type, and the sample accent type is the accurate accent type corresponding to the first speech training samples. For example, these three users may all have a Sichuan accent.
[0052] Such as Figure 2As shown, the initial accent classification model can determine the accent features corresponding to multiple sub-audio segments of speech information, and based on the accent features, the to-be-determined accent type corresponding to the first speech training sample can be determined. For example, the initial accent classification model can perform an averaging process on the accent features corresponding to multiple sub-audio segments of speech information, so as to analyze the overall interpretation characteristics corresponding to the first speech training sample. The server 201 can determine a second loss function according to the difference between the to-be-determined accent type determined by the initial accent classification model and the sample accent type, and adjust the parameters of the initial accent classification model through the second loss function, so that the initial accent classification model can determine a more accurate accent type.
[0053] Among them, since the multiple sub-audio segments of speech information included in the first speech training sample all correspond to users of the same accent type, generally, the multiple sub-audio segments of speech information should have similar accent features. Based on this, during the training process, in order to ensure the rationality of the initial classification model when determining accent features, the server 201 can determine a first loss function according to the difference between the accent features corresponding to multiple sub-audio segments of speech information. By combining the first loss function and the second loss function to adjust the parameters of the initial accent classification model, while enabling the model to perform more accurate accent classification, the model can determine similar accent features for multiple sub-audio segments of speech information corresponding to users of the same accent type, improving the rationality and accuracy of accent feature determination. Thus, the trained accent classification model can not only accurately perform accent classification on the first speech training sample, but also obtain a relatively accurate and reasonable classification result when applied in practice, such as when determining the accent type corresponding to the to-be-classified speech information, improving the practicality of the model.
[0054] Next, in combination with the accompanying drawings, a data processing method provided by an embodiment of the present application will be introduced.
[0055] See Figure 3 , Figure 3 which is a flowchart of a data processing method provided by an embodiment of the present application. The method includes:
[0056] S301: Obtain a first speech training sample corresponding to a sample accent type.
[0057] Among them, the first voice training sample can be a piece of voice information, such as the voice of a user collected, etc. The sample accent type is the accurate accent type corresponding to the first voice training sample. The first voice training sample includes multiple sub-voice information of users with the same accent type, that is, the multiple sub-voice information is the voice information emitted by users with the same accent type. For example, the first voice training sample can be multiple long pieces of voice information spoken by several users with the Sichuan dialect accent, and the multiple sub-voice information can be the voice information corresponding to different time intervals of the voice information.
[0058] It can be understood that according to different requirements for actual accent type classification, the selection of voice training samples can also be different. For example, when it is necessary to train a model that can distinguish the unique accent type of each user, the first voice training sample can be a training sample corresponding to the same user; when it is necessary to train a model that can distinguish the accent types of users in a certain region, the first voice training sample can be a training sample of one or more users corresponding to the same regional accent type, which is not limited here.
[0059] S302: Determine the accent features corresponding to the multiple sub-voice information and the pending accent type corresponding to the first voice training sample according to the first voice training sample and the initial accent classification model.
[0060] The initial accent classification model can be used to determine the accent type corresponding to the voice information, and the first voice training sample can be used to train the initial accent classification model. When determining the accent type corresponding to the first voice training sample, the initial accent classification model can first determine the accent features corresponding to the multiple sub-voice information, and the accent features can reflect the accent characteristics of the users corresponding to the sub-voice information. For example, the multiple sub-voice information corresponding to different time intervals in the first voice training sample can be determined first, and then the accent features corresponding to each sub-voice information can be determined.
[0061] After determining the accent features, the initial accent classification model can determine the accent characteristics of each part of the voice information in the first voice training sample based on the accent features. Since the sub-voice information is a component of the first voice training sample, these accent characteristics can be comprehensively analyzed to determine the pending accent type corresponding to the whole of the first voice training sample. The pending accent type is the classification result determined by the initial accent classification model. For example, the initial accent classification model can determine the average feature value of the accent features of each sub-voice information, and the average feature value can reflect the overall accent characteristics of the first voice training sample, so that the pending accent type corresponding to the first voice training sample can be determined based on the average feature value.
[0062] S303: Determine the first loss function and the second loss function corresponding to the first speech training sample.
[0063] As mentioned above, the sample accent type is the accurate accent type corresponding to the first speech training sample, and the to-be-determined accent type is the accent type determined by the initial accent classification model. Therefore, the difference between the to-be-determined accent type and the sample accent type can reflect the accuracy of the initial accent classification model for accent classification of speech information. To make the accent classification result determined by the initial accent classification model more accurate, the processing device can determine the second loss function corresponding to the first speech training sample according to the difference between the to-be-determined accent type and the sample accent type. The second loss function can be used to enable the initial accent classification model to learn how to determine a more accurate accent classification result for the first speech training sample during the training process.
[0064] It can be understood that when only training the model through the second loss function, although the finally determined accent classification result can be made more accurate, it does not pay attention to the rationality of the way the initial accent classification model determines the accent classification result. As mentioned above, when the model determines the accent classification result, it will first determine the accent features corresponding to each segment of sub-speech information, and determine the to-be-determined accent type of the whole first speech training sample by synthesizing these accent features. When only correcting the final classification result but not paying attention to the intermediate process, it may lead to the overfitting phenomenon of unreasonable parameter adjustment that the initial accent classification model overly pursues the accuracy of the accent classification result of the first speech training sample.
[0065] For example, the accent features determined by the initial accent classification model for each segment of sub-speech information may have large differences, but after comprehensive processing, the determined to-be-determined accent type is very close to the sample accent type. Since multiple segments of sub-speech information correspond to users of the same accent type, and the accent characteristics of users of the same accent type are relatively similar, under normal circumstances, multiple segments of sub-speech information should have relatively similar accent features. Thus, it can be seen that if only the result finally output by the model is constrained, it will lead to unreasonable feature determination results when the model determines the accent features, so that when the model performs accent classification on speech information other than the first speech training sample, it is likely to have a low accuracy of the final accent classification result due to abnormal accent feature determination, and the practicability of the model is poor; at the same time, due to the poor ability of the model to determine accent features, it cannot be applied to other scenarios with requirements for accent features, and the applicability is low.
[0066] Based on this, in the embodiments of the present application, in order to solve the above technical problems, the processing device can also utilize the principle that users with the same accent type have relatively fixed accent characteristics to perform more refined model training on the initial accent classification model. The processing device can determine the first loss function corresponding to the first voice training sample according to the differences between the accent characteristics corresponding to multiple segments of sub-voice information, and this first loss function is used to enable the initial accent classification model to learn how to determine more reasonable and accurate accent characteristics for the multiple segments of sub-voice information of the first voice training sample during the training process.
[0067] S304: Adjust the parameters of the initial accent classification model according to the first loss function and the second loss function to obtain an accent classification model.
[0068] Referring to the roles of the above first loss function and second loss function in the model training process, the processing device can combine the first loss function and the second loss function to adjust the parameters of the initial accent classification model, so that the trained accent classification model can obtain a relatively accurate accent classification result for the first voice training sample, and at the same time, can determine relatively reasonable accent characteristics for the multiple segments of sub-voice information included in the first voice training sample, and control the differences between these accent characteristics within a reasonable range, so as to conform to the accent characteristics of users with the same accent type when speaking. By training the model in this way, the accent classification model can not only be applied to classify the accents of training samples, but also, based on reasonable accent type determination parameters, can be applied to other voice information to be classified, that is, it can be used to determine the accent type corresponding to the voice information to be classified. At the same time, since the accent classification model has relatively accurate determination parameters when determining accent characteristics, it can also be applied to other various scenarios that require accent characteristic determination and has relatively excellent applicability.
[0069] It can be seen from the above technical solutions that adjusting the parameters of the initial accent classification model based on the first loss function and the second loss function can, on the one hand, make the accent type determined by the initial accent classification model more accurate, and on the other hand, can control the differences between the accent characteristics of the determined sub-voice information within a reasonable range during the training process of the model, making the way of determining the accent characteristics more conform to the real accent situation of users with the same accent type, and improving the accuracy and rationality of the training of the accent classification model.
[0070] As mentioned above, through the first loss function and the second loss function, the parameters of the initial accent classification model can be adjusted in terms of both accent characteristic determination and accent type determination to improve the rationality of the obtained accent classification model. Next, the specific parameter adjustment method will be introduced in detail.
[0071] In one possible implementation, the initial accent classification model may include an initial feature extraction sub-model and an initial feature classification sub-model. The initial feature extraction sub-model is used to determine the accent features corresponding to multiple segments of sub-speech information, and the initial feature classification sub-model is used to determine the pending accent type based on the accent features.
[0072] For the initial feature extraction submodel, when adjusting the parameters, in order to make the determined accent features more reasonable and accurate, the processing device can perform parameter features on the initial feature extraction submodel according to the first loss function and the second loss function to obtain the feature extraction submodel. Thus, for the first speech training sample, the feature extraction submodel can make the determined accent features more helpful in determining the accurate accent type on the one hand, and can make the differences between the determined accent features more in line with the accent characteristics of users with the same accent type on the other hand, improve the rationality of the accent features, and achieve two-way improvements on the final classification result and the rationality of the features. For the initial feature classification model, since the model does not involve the process of determining the accent features, the processing device can only use the second loss function to adjust the parameters of the initial feature classification submodel to obtain the feature classification submodel, so that the feature classification submodel can perform more accurate analysis and processing on the accent features, and can determine the pending accent type that is more in line with the accent type of the sample for the first speech training sample. According to the feature extraction submodel and the feature classification submodel, the processing device can determine the accent classification model as the training result of this model training.
[0073] It can be seen that targeted training of different model parts of the initial accent classification model through different loss functions can reasonably and effectively improve the accuracy and rationality of each model part, thereby improving the overall classification effect of the trained accent classification model.
[0074] Among them, in order to reasonably use multiple loss functions to adjust the parameters of the same sub-model, in a possible implementation, the processing device can configure corresponding weight parameters for the first loss function and the second loss function based on the experience in past model training or the parameter settings of relevant personnel. The processing device can determine the first weight parameter corresponding to the first loss function and the second weight parameter corresponding to the second loss function, the first weight parameter is used to identify the influence of the first loss function on the adjustment of the parameters of the initial feature extraction sub-model, and the second weight parameter is used to identify the influence of the second loss function on the adjustment of the parameters of the initial feature extraction sub-model.
[0075] When performing model training, the processing device can determine a comprehensive loss function based on the first loss function, the first weight parameter, the second loss function, and the second weight parameter, and adjust the parameters of the initial feature extraction sub-model according to the comprehensive loss function to obtain a feature extraction sub-model. Thus, the influence degree of different loss functions on parameter adjustment can be reasonably controlled through the weight parameter, enabling the finally obtained model to perform feature extraction more accurately. The extracted accent features can not only conform to the characteristics of the real user accent but also help determine the accent type corresponding to the voice information.
[0076] In addition to the diverse ways of parameter adjustment, the specific model types of each sub-model in the initial accent classification model can also include multiple types. For example, in a possible implementation manner, to further improve the efficiency of model training and reduce the manual participation required for model training, the processing device can obtain the initial feature extraction sub-model through self-supervised learning.
[0077] Among them, self-supervised learning means that during the training process of the model, the model can extract some data from the sample data as labels by itself, and analyze the accuracy of the model output results through these labels to perform corresponding parameter adjustments. In the embodiments of the present application, before determining the accent features corresponding to multiple segments of sub-voice information based on the first voice training sample and the initial accent classification model, the processing device can first obtain a self-supervised learning model and a second voice training sample, and the second voice training sample is used to train the self-supervised learning model.
[0078] The processing device can first determine the target voice information corresponding to the second voice training sample through the self-supervised learning model. The target voice information corresponds to the target voice part in the second voice training sample, that is, the self-supervised learning model can extract the voice information of the target voice part in the second voice training sample as a training label. Based on different label extraction methods, the target voice part may also be different, which is not limited here.
[0079] To enable the self-supervised learning model to learn the method of extracting speech information features, the processing device can use the self-supervised learning model to determine the pending speech information corresponding to the target speech part based on the second speech training sample after removing the target speech information. That is, the processing device can enable the self-supervised learning model to perform feature analysis on the remaining speech information in the second speech training sample after removing the target speech information during the training process, and predict the missing speech information of the target speech part. The processing device can determine the third loss function corresponding to the third speech training sample according to the difference between the target speech information and the pending speech information. Since the target speech information is the true speech information corresponding to the target speech part, and the pending speech information is the speech information of the target speech part determined by the self-supervised learning model through feature analysis of the remaining speech information part of the second speech training sample, this difference can reflect the accuracy of the self-supervised learning model in analyzing the features of speech information.
[0080] Therefore, the processing device can adjust the parameters of the self-supervised learning model according to the third loss function, enabling the self-supervised learning model to learn how to more accurately analyze the features of speech information. After parameter adjustment, since the obtained model has a certain feature analysis ability, the initial feature extraction sub-model can be determined based on the adjusted self-supervised learning model to determine the accent features corresponding to the speech information. Throughout the training process, there is no need for manual annotation of the labels corresponding to the samples. The self-supervised learning model itself only needs to determine part of the speech information in the sample as the label, and then predict the missing part of the speech information based on the features of the remaining speech information, thereby learning how to accurately analyze the features of speech information, reducing the demand for human resources, and improving the convenience of model training.
[0081] As Figure 4 shown, Figure 4 Fig. is a schematic diagram of model training. The input of the model is the original speech waveform (Raw waveform) X, and the model mainly includes the following three parts:
[0082] 1. Feature encoder: This part can be composed of a multi-layer convolutional neural network (Convolutional Neural Networks, CNN), which is responsible for extracting the implicit speech features (Latent speech representations) Z from the original speech waveform.
[0083] 2. Context network: It is mainly composed of multiple layers of Transformer structures. The output Z of the feature encoder part passes through the activation function (Gaussian Error Linear Units, abbreviated as GELU) layer and can be used as the input of this part. Finally, the context representation C is output. This context representation is the speech information of the masked part determined by the model through the input speech information.
[0084] 3. Quantization module: Discretizes the Z output by the feature encoder to obtain the quantized representation Q. Q can be used as a label (i.e., supervision information) to help train this model.
[0085] In the actual training process, first, a masking operation is performed on Z, that is, the information part corresponding to Q in Z is covered, and then it is input into the context network. This context network can determine the information in the masked interval based on the remaining speech information in Z. The output at time t of the masked interval can be c t , corresponding to q in the quantized representation after quantization t , Q t represents the set of all candidate quantized representations, including q t and k interference parameters. The loss function of this model can be expressed as:
[0086]
[0087] where sim(a, b) = a T b / ∥a∥∥b∥, representing the correlation between two vectors. This loss function can adjust the parameters of this model based on the difference between Q and C, so that this model can analyze the features of speech information more accurately.
[0088] As Figure 5 shown, Figure 5A schematic diagram of an initial accent classification model provided by an embodiment of the present application. After pre-training the self-supervised learning model with a large amount of unlabeled speech data, the self-supervised learning model already has a certain ability to accurately extract features in speech information. The processing device can extract the feature encoder and the context network layer therein as the initial feature extraction sub-model in the initial accent classification model. The processing device can add a linear layer (Affine) to this sub-model, and this linear layer can be used to linearly classify the accent features determined by this sub-model based on relevant parameters. For example, it can determine the probabilities of various accents corresponding to the speech information based on the accent features. After the results of this linear classification are averaged (Mean) and normalized (Softmax), the pending accent type of the input speech information can be determined. This linear layer, the average layer, and the normalization layer can form the initial feature classification sub-model in this model. The processing device can determine the cross-entropy loss function L based on the difference between this pending accent type and the sample accent type CE , as the second loss function for parameter adjustment.
[0089] At the same time, since this linear layer linearly classifies the accent features of multiple segments of sub-speech information output by this initial feature classification sub-model with the same parameter, the differences between the linear classification results corresponding to multiple segments of sub-speech information can reflect the differences between the accent features. Through this difference, the first loss function can be determined for parameter adjustment. For example, in Figure 5 In the model structure shown, the multiple segments of sub-speech information can be speech information divided according to speech frames. Assuming there are N accent types in total, a 768*N linear layer can be added on the basis of the context network, and the parameters of this linear layer are randomly initialized. Assuming the output of this linear layer for the i-th frame of sub-speech information is A i , the processing device can average the A corresponding to multiple frames of sub-speech information, and after normalization, obtain the pending accent type. At the same time, the standard deviation of the A corresponding to multiple frames of sub-speech information is calculated in the time dimension to obtain i , and according to this standard deviation, the standard deviation constraint loss function L i can be determined. As shown in the following formula: , where N is the total number of frames of the input speech information. This standard deviation constraint loss function can reflect the differences between the accent features of multiple frames of sub-speech information. Therefore, this standard deviation constraint loss function can be used as the first loss function in the embodiment of the present application. Thus, the loss function L corresponding to this initial accent classification model can be expressed as the sum of this cross-entropy loss function and the standard deviation constraint loss function: SDC As shown in the following formula:
[0090]
[0091] N is the total number of frames of the input speech information. This standard deviation constraint loss function can reflect the differences between the accent features of multiple frames of sub-speech information. Therefore, this standard deviation constraint loss function can be used as the first loss function in the embodiment of the present application. Thus, the loss function L corresponding to this initial accent classification model can be expressed as the sum of this cross-entropy loss function and the standard deviation constraint loss function:
[0092] L=L SDC +L CE
[0093] The loss function L can be used to adjust the parameters of the initial feature extraction submodule. CE It can be used to adjust the parameters of the linear layer in the initial feature classification submodule.
[0094] It can be understood that in the above model training process, the second loss function is used for the parameter adjustment of the initial feature extraction submodel and the initial feature classification submodel. In the initial feature extraction submodel, the second loss function can help to make the determined accent features more helpful in determining the accurate accent type, and in the initial feature classification submodel, the second loss function can help to make the classification results of the accent features of the submodel more reasonable, while the first loss function only constrains the parameter adjustment of the initial feature extraction submodel. It can be seen that compared with the first loss function, the second loss function has a greater impact on the overall model training process.
[0095] Based on this, in order to further improve the efficiency and accuracy of model training, in one possible implementation method, before combining multiple loss functions to adjust parameters, the processing device can first reduce the parameter adjustment range corresponding to the second loss function as much as possible, that is, improve the adjustment accuracy of the second loss function during parameter adjustment, thereby improving the accuracy of the overall parameter adjustment of the model.
[0096] Among them, since the second loss function is determined based on the difference between the pending accent type and the sample accent type, and the pending accent type is mainly classified based on the initial feature classification sub-model, therefore, in order to improve the adjustment accuracy of the second loss function, the processing device may first train the initial feature classification sub-model for a preset number of times during the training of the initial accent classification model, and then combine the first loss function and the second loss function to perform comprehensive parameter adjustments on the initial feature extraction sub-model and the initial feature classification sub-model. The preset number of times may be determined by the processing device based on past model training data, or may be set by relevant personnel based on model training experience, and is not limited here.
[0097] After determining the first loss function and the second loss function based on the first voice training sample, the processing device may first determine whether the number of parameter adjustments to the initial feature classification sub-model has reached a preset number. In response to the number of parameter adjustments to the initial feature classification sub-model reaching the preset number, the processing device may adjust the parameters of the initial feature extraction sub-model according to the first loss function and the second loss function to obtain the feature extraction sub-model. Correspondingly, if the processing device determines that the preset number has not been reached, it may only perform parameter adjustment on the initial feature classification sub-model based on the second loss function, thereby further improving the parameter adjustment accuracy corresponding to the second loss function and further improving the overall training efficiency of model training.
[0098] Meanwhile, in this model training scenario, the initial feature extraction sub-model only depends on its own model parameters to determine the accent features, and there is no problem of accuracy of the required voice samples. Therefore, the initial feature extraction sub-model can undergo a certain pre-training process and usually already has a certain model accuracy. For example, when the initial feature extraction sub-model is trained through a self-supervised learning model, the initial feature extraction sub-model already has a certain feature extraction ability for voice information; for the initial feature classification sub-model, the parameter adjustment of this sub-model requires the output features of the initial feature extraction sub-model. Therefore, for the initial accent classification model, it is very difficult for the initial feature classification sub-model to perform corresponding pre-training. Therefore, the accuracy of the initial feature classification sub-model is usually lower than that of the initial feature extraction sub-model.
[0099] On this premise, by first performing parameter adjustment on the initial feature classification sub-model for a preset number of times and then combining multiple loss functions to adjust the parameters of the overall model, the negative impact of the low-precision model on the high-precision model can be reduced to a certain extent, so that the parameter adjustment in the entire model training process tends to be adjusted in a positive direction that is conducive to improving the model accuracy, thereby improving the efficiency of model training.
[0100] For example, after retaining the scene network and the feature encoder in the self-supervised learning model, the processing device may add a randomly initialized linear layer to the scene network. The parameters of the scene network and the feature encoder already have a certain feature extraction accuracy after pre-training through self-supervised learning, while the feature classification accuracy of the initialized linear layer is relatively low. Therefore, in the early stage of training, the processing device may first keep the model parameters in the scene network and the feature encoder unchanged and only update the parameters in the linear layer, so that 2000 iterations can be performed to first improve the classification accuracy of the linear layer; in the later stage of training, the overall parameters of the model are adjusted based on multiple loss functions.
[0101] In addition, in a possible implementation, to make the initial feature extraction sub-model trained by the above method more suitable for feature extraction in the initial accent classification model, the processing device can adjust the parameters of the initial accent classification model and the self-supervised learning model using the same training samples. For example, the first speech training sample can be the same as the second speech training sample, thereby reducing the negative impact brought by different training samples during the model training process. For example, it can reduce the possibility that the initial accent classification model and the self-supervised learning model learn conflicting parameter adjustment directions. While ensuring the accuracy of model training, it can also improve the efficiency of model training to a certain extent and reduce the amount of sample data required for training.
[0102] It can be understood that in real life, due to the wide range of regions where users are located and the complex geographical environment in which they live, there are multiple accent types. Based on this, in addition to using consistent training samples, in a possible implementation, to make the trained model more widely applicable in real life and be able to accurately identify and classify speech information of multiple accent types, the processing device can obtain a speech training sample set, and the first speech training sample is any one of the multiple first speech training samples included in the speech training sample set. That is, each first speech training sample in the speech training sample set includes multiple sub-speech information of users corresponding to the same accent type. At the same time, to improve the model's recognition ability for multiple accent types, the multiple sub-speech information included in the multiple first speech training samples can correspond to users of multiple accent types, such as the first user and the second user, and the accent type of the first user is different from that of the second user. Of course, on the basis of ensuring the diversity of accent types, there can also be multiple users of one accent type, and only the smallest component is introduced here. Thus, based on the speech training sample set, the initial accent classification model can be trained with the speech information of users of multiple accent types to improve the model training effect and the applicability of the model in actual applications.
[0103] Among them, in order to highlight the influence of different accents on speech information and reduce the interference of other interfering factors on the accent recognition ability of the accent classification model, in a possible implementation manner, the multiple first speech training samples may correspond to the same language. Thus, the speech training samples corresponding to different accent types actually only have differences in accents and no differences in languages. During the training process, the model can more focusedly learn the distinguishing features between different accents and achieve a more accurate classification of accent types. At the same time, in order to make the accent types of different users more clearly distinguishable and improve the effectiveness of model training, the processing device may define that users with different accent types correspond to different accent regions, and the accent types corresponding to different accent regions are different, that is, each user in an accent region has its own unique accent type. For example, when the same language is Chinese, the different accent regions may be different countries. For example, when British people, Chinese people, and Japanese people speak Chinese, they may all have the accent characteristics of their native languages, so that speech training samples of Chinese with different accent types can be collected.
[0104] To facilitate understanding of the technical solution provided by the embodiments of the present application, next, a data processing method provided by the embodiments of the present application will be introduced in combination with an actual application scenario.
[0105] See Figure 6 , Figure 6 which is a flowchart of a data processing method in an actual application scenario provided by the embodiments of the present application. In this actual application scenario, the processing device may be a server for model training.
[0106] The method includes:
[0107] S601: Obtain a speech training sample set including multiple first speech training samples.
[0108] Among them, each first speech training sample corresponds to a user. The users corresponding to the multiple sub-speech information in different first speech training samples may be different, and different users have different accent characteristics. In this actual application scenario, the collected speech information may all be English speech, and different users have different national backgrounds, so that English speech with different accent types can be obtained, such as Chinese English accent, American English accent, British English accent, etc.
[0109] S602: Train a self-supervised learning model according to the speech training sample set to obtain an initial feature extraction sub-model.
[0110] Here, the speech training samples excluding the sample accent types in the speech training sample set can be used for training, and the parameter settings of the model are as follows:
[0111] Feature Encoder: A 7-layer CNN network is used, with each CNN network layer having 512 channels, corresponding strides of (5, 2, 2, 2, 2, 2, 2), and corresponding convolutional kernel sizes of (10, 3, 3, 3, 3, 2, 2).
[0112] Scene Network: A 12-layer compiler is used, the model dimension is 768, the internal fully connected layer dimension is 3072, and 8 heads are used for multi-head attention.
[0113] During the training process, Adam can be used as the optimization function. A total of 400K iterations are trained. The first 8% of the iterations use the warm-up learning rate, with a maximum learning rate of 0.005, and the learning rate of the subsequent iterations decreases linearly.
[0114] S603: Determine whether the accent classification accuracy has reached a preset value.
[0115] Through the feature encoder and scene network in the trained self-supervised learning model, and adding a linear layer for feature classification, the server can obtain an initial accent classification model. During each training iteration, the server can determine whether the accent classification accuracy of this model has reached the preset value. If not, it can jump to S604. If so, it can jump to S608 to obtain the accent classification model and complete the training. In addition, the server can also determine whether parameter adjustment has been performed on each training sample in the speech training sample set, and make the model use all the training samples during the training process to improve the comprehensiveness of model training.
[0116] S604: Obtain a first speech training sample from the speech training sample set, and determine the first loss function and the second loss function corresponding to the first speech training sample.
[0117] The first speech training sample can be an unused training sample or a used training sample for repeated training, which is not limited here. The server can input the first speech training sample into the initial accent classification model to obtain the accent features corresponding to multiple sub-speech information segments therein, and the pending accent type corresponding to the first speech training sample. The first loss function can be determined according to the differences between the accent features, and the second loss function can be determined according to the differences between the pending accent type and the sample accent type.
[0118] S605: Determine whether the number of parameter adjustment times for the initial feature classification sub-model has reached the preset number of times.
[0119] During this training process, the preset number of times can be 2000 times. In the first 2000 iterations, the parameters of the initial feature extraction sub-model remain unchanged, and only the model parameters of the initial feature classification sub-model are updated. Starting from the 2000th iteration, all model parameters are continuously adjusted, and the initial learning rate can be set to 0.00002.
[0120] S606: Adjust the parameters of the initial feature classification sub-model according to the second loss function.
[0121] If the preset number of times is not reached, it means that the accuracy of the initial feature classification sub-model needs to be further improved. The server can only adjust the parameters of the initial feature classification sub-model based on the second loss function.
[0122] S607: Adjust the parameters of the initial feature extraction sub-model according to the first loss function and the second loss function, and adjust the parameters of the initial feature classification sub-model according to the second loss function.
[0123] If the preset number of times is reached, it means that the accuracy of the initial feature sub-classification model can already be used for the overall model training process.
[0124] After parameter adjustment, the server can re-execute the operation of S603 to determine whether the model training is completed.
[0125] To compare with the related technologies, the server can test on the same speech training sample set through multiple accent classification models in the related technologies. For example, it can include the following three models:
[0126] (1) A multi-step model using i-vector feature vectors
[0127] As introduced before, the common technical solutions include multi-step solutions and end-to-end solutions. The multi-step model generally first extracts feature vectors from the speech data, such as i-vector and x-vector, and then connects a classifier for final classification. For comparison, two traditional methods are used here for accent classification, namely the accent classifier based on i-vector and the accent classifier based on x-vector. The input of both models is 23-dimensional speech features, including 20-dimensional Mel Frequency Cepstral Coefficients (MFCC for short) and 3-dimensional pitch parameters.
[0128] The processing flow of the accent classifier based on i-vector includes:
[0129] 1. Train the i-vector model.
[0130] 2. Extract 600-dimensional i-vector features from the speech data.
[0131] 3. Post-process the i-vector features, including mean normalization, LDA dimensionality reduction to 50 dimensions, length regularization, etc.
[0132] 4. Train a logistic regression classifier for classification.
[0133] (2) Multi-step model using x-vector feature vectors
[0134] 1. Train an x-vector feature extractor.
[0135] 2. Extract 512-dimensional x-vector features from the speech data.
[0136] 3. Post-process the x-vector features, including mean normalization, LDA dimensionality reduction to 50 dimensions, length regularization, etc.
[0137] 4. Train a logistic regression classifier for classification.
[0138] (3) End-to-end model with only cross-entropy loss function
[0139] The end-to-end solution generally uses a pooling layer (Pooling) to convert the frame-level vectors output by the network into sentence-level vectors, then uses a linear layer for classification, and finally uses the cross-entropy loss function to train the entire network uniformly. This model can be as Figure 7 shown. The specific model training process includes:
[0140] 1. Pre-train to obtain a self-supervised learning model. It should be emphasized that the end-to-end models in the related technologies do not use self-supervised learning models for accent feature extraction. This solution is an improvement of the end-to-end models in the related technologies to emphasize the role of the first loss function in model training.
[0141] 2. Similar to the technical solution of this application, the context network and feature encoder in the self-supervised learning model can be used for feature extraction, and a sentence-level pooling layer is added to the context network. This pooling layer includes taking the average and standard deviation, and then the obtained vectors are concatenated.
[0142] 3. Add a linear layer on the pooling layer for accent classification.
[0143] 4. Obtain the cross-entropy loss function according to the to-be-determined accent type and the sample accent type for model training.
[0144] During the above training process, the self-supervised learning model can also be trained using the 960-hour speech corpus (LibriSpeech) data. The speech training sample set used can be English with 8 accents, namely Russian, Korean, American, Portuguese, Japanese, Indian, British, and Chinese, with approximately 20 hours of training data for each accent. The specific content is as follows:
[0145]
[0146]
[0147] The training results are shown in the following table:
[0148] Accent x-vector i-vector AI0 AI1 AM 46.4 48.8 87.2 89.6 BR 66.0 72.1 67.7 83.2 CH 62.2 70.9 61.5 54.0 IN 77.4 87.4 63.0 65.2 JA 53.5 45.9 94.8 95.8 KO 56.1 58.1 74.7 70.2 PO 54.7 58.0 82.0 82.7 RU 59.9 64.3 51.4 52.6 All accents 59.1 62.6 72.7 73.9
[0149] Among them, AI0 is an end-to-end model trained only using the cross-entropy loss function, and AI1 is an end-to-end model in this application that combines the standard deviation constraint loss function that can reflect the differences between accent features. As can be seen from the table, the solution provided in this application is superior to the solutions in the related art in terms of the classification accuracy of the vast majority of accent types, showing a significant improvement.
[0150] The accent classification model trained by this method can be applied to various scenarios. For example, it can include the three major scenarios such as Figure 8 shown, namely intelligent speech, intelligent map, and intelligent audio-visual. Specifically, it can be divided into small scenarios such as commuting, traveling, and voice interaction. The speech recognition service can be used for vehicle networking voice interaction, as well as other intelligent hardware such as speakers and robots.
[0151] Based on the data processing method provided in the above embodiments, the embodiments of the present application also provide a data processing device. Refer to Figure 9 , Figure 9 which is the structural block diagram of a data processing device 900 provided in the embodiments of the present application. The device 900 includes a first acquisition unit 901, a first determination unit 902, a second determination unit 903, and a first adjustment unit 904:
[0152] The first acquisition unit 901 is configured to acquire a first speech training sample corresponding to a sample accent type, and the first speech training sample includes multiple sub-speech information of users with the same accent type;
[0153] The first determination unit 902 is configured to determine the accent features corresponding to the multiple sub-speech information and the pending accent type corresponding to the first speech training sample according to the first speech training sample and the initial accent classification model, and the pending accent type is determined based on the accent features;
[0154] A second determination unit 903, configured to determine a first loss function and a second loss function corresponding to the first speech training sample, where the first loss function is determined according to the differences between the accent features corresponding to the multiple sub-speech information segments, and the second loss function is determined according to the difference between the to-be-determined accent type and the sample accent type;
[0155] A first adjustment unit 904, configured to adjust parameters of the initial accent classification model according to the first loss function and the second loss function, to obtain an accent classification model, where the accent classification model is used to determine the accent type corresponding to the speech information to be classified.
[0156] In a possible implementation manner, the initial accent classification model includes an initial feature extraction sub-model and an initial feature classification sub-model. The initial feature extraction sub-model is configured to determine the accent features corresponding to the multiple sub-speech information segments respectively, and the initial feature classification sub-model is configured to determine the to-be-determined accent type according to the accent features. The first adjustment unit 904 is specifically configured to:
[0157] Adjust parameters of the initial feature extraction sub-model according to the first loss function and the second loss function, to obtain a feature extraction sub-model;
[0158] Adjust parameters of the initial feature classification sub-model according to the second loss function, to obtain a feature classification sub-model;
[0159] Determine the accent classification model according to the feature extraction sub-model and the feature classification sub-model.
[0160] In a possible implementation manner, the apparatus 900 further includes a third determination unit:
[0161] The third determination unit is configured to determine a first weight parameter corresponding to the first loss function and a second weight parameter corresponding to the second loss function;
[0162] The first adjustment unit 904 is specifically configured to:
[0163] Determine a comprehensive loss function according to the first loss function, the first weight parameter, the second loss function, and the second weight parameter;
[0164] Adjust parameters of the initial feature extraction sub-model according to the comprehensive loss function, to obtain a feature extraction sub-model.
[0165] In a possible implementation manner, the apparatus 900 further includes a second acquisition unit, a fourth determination unit, a fifth determination unit, a sixth determination unit, and a second adjustment unit:
[0166] A second acquisition unit, configured to acquire a self-supervised learning model and second speech training samples;
[0167] A fourth determination unit, configured to determine, by using the self-supervised learning model, target speech information corresponding to the second speech training samples, where the target speech information corresponds to a target speech part in the second speech training samples;
[0168] A fifth determination unit, configured to determine, by using the self-supervised learning model and according to the second speech training samples after removing the target speech information, pending speech information corresponding to the target speech part;
[0169] A sixth determination unit, configured to determine a third loss function corresponding to the second speech training samples according to a difference between the target speech information and the pending speech information;
[0170] A second adjustment unit, configured to adjust parameters of the self-supervised learning model according to the third loss function, and determine the initial feature extraction sub-model according to the adjusted self-supervised learning model.
[0171] In a possible implementation manner, the first adjustment unit 904 is specifically configured to:
[0172] In response to that the number of parameter adjustment times for the initial feature classification sub-model reaches a preset number of times, adjust parameters of the initial feature extraction sub-model according to the first loss function and the second loss function, to obtain the feature extraction sub-model.
[0173] In a possible implementation manner, the first speech training samples are the same as the second speech training samples.
[0174] In a possible implementation manner, the first speech training sample is any one of multiple first speech training samples included in a speech training sample set, where multiple segments of sub-speech information included in the multiple first speech training samples correspond to a first user and a second user, and an accent type of the first user is different from an accent type of the second user.
[0175] In a possible implementation manner, the multiple first speech training samples correspond to the same language, and users with different accent types correspond to different accent regions.
[0176] An embodiment of this application further provides a computer device, which is introduced below with reference to the accompanying drawings. Please refer to Figure 10As shown, an embodiment of the present application provides a device, which may also be a terminal device. The terminal device may be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA for short), a point of sales (POS for short), an in-vehicle computer, etc. Taking the terminal device as a mobile phone as an example:
[0177] Figure 10 The block diagram of a part of the structure of the mobile phone related to the terminal device provided by the embodiment of the present application is shown. Refer to Figure 10 , the mobile phone includes: a radio frequency (RF) circuit 710, a memory 720, an input unit 730, a display unit 740, a sensor 750, an audio circuit 760, a wireless fidelity (WiFi) module 770, a processor 780, and a power supply 790 and other components. Those skilled in the art can understand that Figure 10 the structure of the mobile phone shown in
[0178] does not constitute a limitation on the mobile phone, and may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements. Figure 10 The following specifically introduces each component of the mobile phone in combination with
[0179] The RF circuit 710 can be used for receiving and sending signals during information reception or call processes. Specifically, after receiving the downlink information of the base station, it is given to the processor 780 for processing; in addition, the designed uplink data is sent to the base station. Generally, the RF circuit 710 includes but is not limited to antennas, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA for short), a duplexer, etc. In addition, the RF circuit 710 can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the global system of mobile communication (GSM for short), general packet radio service (GPRS for short), code division multiple access (CDMA for short), wideband code division multiple access (WCDMA for short), long term evolution (LTE for short), email, short messaging service (SMS for short), etc.
[0180] The memory 720 can be used to store software programs and modules. The processor 780 executes various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 720. The memory 720 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 720 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0181] The input unit 730 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the mobile phone. Specifically, the input unit 730 may include a touch panel 731 and other input devices 732. The touch panel 731, also known as a touch screen, can collect touch operations of the user on or near it (such as operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch panel 731), and drive corresponding connection devices according to a preset program. Optionally, the touch panel 731 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 780, and can receive and execute commands sent by the processor 780. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch panel 731. In addition to the touch panel 731, the input unit 730 may further include other input devices 732. Specifically, the other input devices 732 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.
[0182] The display unit 740 can be used to display information input by the user or information provided to the user, as well as various menus of the mobile phone. The display unit 740 may include a display panel 741. Optionally, the display panel 741 can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, the touch panel 731 can cover the display panel 741. When the touch panel 731 detects a touch operation on or near it, it is transmitted to the processor 780 to determine the type of touch event. Subsequently, the processor 780 provides a corresponding visual output on the display panel 741 according to the type of touch event. Although in Figure 10 the touch panel 731 and the display panel 741 are implemented as two independent components to realize the input and input functions of the mobile phone, in some embodiments, the touch panel 731 and the display panel 741 can be integrated to realize the input and output functions of the mobile phone.
[0183] The mobile phone may further include at least one sensor 750, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 741 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 741 and / or the backlight when the mobile phone is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity, and can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors such as gyroscope, barometer, hygrometer, thermometer, infrared sensor that the mobile phone can also be configured with, they will not be elaborated here.
[0184] The audio circuit 760, the speaker 761, and the microphone 762 can provide an audio interface between the user and the mobile phone. The audio circuit 760 can transmit the electrical signal converted from the received audio data to the speaker 761, and the speaker 761 converts it into a sound signal for output; on the other hand, the microphone 762 converts the collected sound signal into an electrical signal, which is received by the audio circuit 760 and then converted into audio data. After the audio data is output to the processor 780 for processing, it is sent to another mobile phone through the RF circuit 710, for example, or the audio data is output to the memory 720 for further processing.
[0185] WiFi belongs to short-range wireless transmission technology. The mobile phone can help users send and receive emails, browse the web, and access streaming media through the WiFi module 770. It provides users with wireless broadband Internet access. Although Figure 10The WiFi module 770 is shown, but it can be understood that it does not belong to the essential components of the mobile phone and can be completely omitted within the scope of not changing the essence of the invention as needed.
[0186] The processor 780 is the control center of the mobile phone, connecting various parts of the entire mobile phone through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 720, and by calling data stored in the memory 720, it executes various functions of the mobile phone and processes data. Optionally, the processor 780 may include one or more processing units; preferably, the processor 780 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 780 either.
[0187] The mobile phone also includes a power supply 790 (such as a battery) for supplying power to each component. Preferably, the power supply can be logically connected to the processor 780 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system.
[0188] Although not shown, the mobile phone may also include a camera, a Bluetooth module, etc., which will not be elaborated here.
[0189] In this embodiment, the processor 780 included in the terminal device further has the following functions:
[0190] Obtain a first voice training sample corresponding to a sample accent type, where the first voice training sample includes multiple sub-voice messages of users with the same accent type;
[0191] Determine the accent features corresponding to the multiple sub-voice messages respectively according to the first voice training sample and the initial accent classification model, and the pending accent type corresponding to the first voice training sample, where the pending accent type is determined based on the accent features;
[0192] Determine a first loss function and a second loss function corresponding to the first voice training sample, where the first loss function is determined according to the differences between the accent features corresponding to the multiple sub-voice messages respectively, and the second loss function is determined according to the difference between the pending accent type and the sample accent type;
[0193] Adjust the parameters of the initial accent classification model according to the first loss function and the second loss function to obtain an accent classification model, where the accent classification model is used to determine the accent type corresponding to the voice information to be classified.
[0194] The embodiment of the present application also provides a server. Please refer to Figure 11As shown Figure 11 FIG. 800 is a structural diagram of a server 800 provided by an embodiment of the present application. The server 800 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 822 (for example, one or more processors) and a memory 832, and one or more storage media 830 (for example, one or more mass storage devices) for storing application programs 842 or data 844. Among them, the memory 832 and the storage media 830 may be transient storage or persistent storage. The program stored in the storage media 830 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 822 may be configured to communicate with the storage media 830 and execute a series of instruction operations in the storage media 830 on the server 800.
[0195] The server 800 may further include one or more power supplies 826, one or more wired or wireless network interfaces 850, one or more input / output interfaces 858, and / or one or more operating systems 841, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.
[0196] The steps performed by the server in the above embodiments may be based on Figure 11 the server structure shown.
[0197] An embodiment of the present application further provides a computer-readable storage medium for storing a computer program, and the computer program is used to execute any one of the data processing methods described in the foregoing embodiments.
[0198] An embodiment of the present application further provides a computer program product including instructions, and when the computer program product runs on a computer, the computer is caused to execute the data processing method provided in any one of the above embodiments.
[0199] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program may be stored in a computer-readable storage medium, and when the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium may be at least one of the following media: read-only memory (abbreviation: ROM), RAM, magnetic disk, or optical disc, etc., which can store program codes.
[0200] It should be noted that the various embodiments in this specification are described in a progressive manner. For the parts that are the same or similar among the various embodiments, reference can be made to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the corresponding parts of the method embodiments. The device and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0201] As described above, it is only a specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in this application should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A data processing method, characterized in that, The method includes: Obtaining a first voice training sample corresponding to a sample accent type, where the first voice training sample includes multiple sub-voice messages of users with the same accent type. The first voice training sample is any one of multiple first voice training samples included in a voice training sample set. The multiple sub-voice messages included in the multiple first voice training samples correspond to a first user and a second user, and the accent type of the first user is different from that of the second user; Determining the accent features corresponding to the multiple sub-voice messages respectively, and the pending accent type corresponding to the first voice training sample according to the first voice training sample and an initial accent classification model. The pending accent type is determined based on the accent features. The initial accent classification model is used to classify the accent type of voice messages, and the accent features reflect the accent characteristics of the users corresponding to the sub-voice messages; Determining a first loss function and a second loss function corresponding to the first voice training sample. The first loss function is determined according to the differences between the accent features corresponding to the multiple sub-voice messages respectively, and the second loss function is determined according to the difference between the pending accent type and the sample accent type; Adjusting the parameters of the initial accent classification model according to the first loss function and the second loss function to obtain an accent classification model, where the accent classification model is used to determine the accent type corresponding to a voice message to be classified; Determining a first weight parameter corresponding to the first loss function and a second weight parameter corresponding to the second loss function. The first weight parameter is used to identify the influence degree of the first loss function on the parameter adjustment of the initial feature extraction sub-model, and the second weight parameter is used to identify the influence degree of the second loss function on the parameter adjustment of the initial feature extraction sub-model; Wherein, the initial accent classification model includes an initial feature extraction sub-model and an initial feature classification sub-model. The initial feature extraction sub-model is used to determine the accent features corresponding to the multiple sub-voice messages respectively, and the initial feature classification sub-model is used to determine the pending accent type according to the accent features. The initial feature extraction sub-model includes a linear layer, and the linear layer is used to perform linear classification on the accent features based on relevant parameters. Adjusting the parameters of the initial accent classification model according to the first loss function and the second loss function to obtain an accent classification model includes: Adjusting the parameters of the initial feature extraction sub-model according to the first loss function and the second loss function to obtain a feature extraction sub-model, where it includes: determining a comprehensive loss function according to the first loss function, the first weight parameter, the second loss function, and the second weight parameter; adjusting the parameters of the linear layer of the initial feature extraction sub-model according to the comprehensive loss function to obtain a feature extraction sub-model; Adjusting the parameters of the initial feature classification sub-model according to the second loss function to obtain a feature classification sub-model; Determining the accent classification model according to the feature extraction sub-model and the feature classification sub-model.
2. The method according to claim 1, wherein Before determining the accent features corresponding to the multi-segment sub-speech information according to the first speech training sample and the initial accent classification model, the method further includes: Obtaining a self-supervised learning model and a second speech training sample; Determining, by the self-supervised learning model, target speech information corresponding to the second speech training sample, where the target speech information corresponds to a target speech part in the second speech training sample; Determining, by the self-supervised learning model, pending speech information corresponding to the target speech part according to the second speech training sample after removing the target speech information; Determining a third loss function corresponding to the second speech training sample according to the difference between the target speech information and the pending speech information; Adjusting the parameters of the self-supervised learning model according to the third loss function, and determining the initial feature extraction sub-model according to the adjusted self-supervised learning model.
3. The method according to claim 1, wherein The adjusting the parameters of the initial feature extraction sub-model according to the first loss function and the second loss function to obtain a feature extraction sub-model includes: In response to the number of parameter adjustment times for the initial feature classification sub-model reaching a preset number of times, adjusting the parameters of the initial feature extraction sub-model according to the first loss function and the second loss function to obtain the feature extraction sub-model.
4. The method according to claim 2, characterized in that The first speech training sample is the same as the second speech training sample.
5. The method according to claim 1, wherein The multiple first speech training samples correspond to the same language, and users with different accent types correspond to different accent regions.
6. A data processing device, characterized in that, The apparatus includes a first acquisition unit, a first determination unit, a second determination unit, a third determination unit, and a first adjustment unit: The first acquisition unit is configured to acquire a first speech training sample corresponding to a sample accent type, where the first speech training sample includes multi-segment sub-speech information of users corresponding to the same accent type, the first speech training sample is any one of the multiple first speech training samples included in the speech training sample set, the multi-segment sub-speech information included in the multiple first speech training samples corresponds to a first user and a second user, and the accent type of the first user is different from the accent type of the second user; The first determination unit is configured to determine the accent features corresponding to the multi-segment sub-speech information respectively according to the first speech training sample and the initial accent classification model, and a pending accent type corresponding to the first speech training sample, where the pending accent type is determined based on the accent features, the initial accent classification model is used to classify the accent type of speech information, and the accent features reflect the accent characteristics of the user corresponding to the sub-speech information; The second determination unit is configured to determine a first loss function and a second loss function corresponding to the first speech training sample, where the first loss function is determined according to the difference between the accent features corresponding to the multi-segment sub-speech information respectively, and the second loss function is determined according to the difference between the pending accent type and the sample accent type; The first adjustment unit is configured to adjust the parameters of the initial accent classification model according to the first loss function and the second loss function to obtain an accent classification model, which is used to determine the accent type corresponding to the speech information to be classified; The third determination unit is configured to determine a first weight parameter corresponding to the first loss function and a second weight parameter corresponding to the second loss function. The first weight parameter is used to identify the influence degree of the first loss function on the parameter adjustment of the initial feature extraction sub-model, and the second weight parameter is used to identify the influence degree of the second loss function on the parameter adjustment of the initial feature extraction sub-model; Wherein, the initial accent classification model includes an initial feature extraction sub-model and an initial feature classification sub-model. The initial feature extraction sub-model is configured to determine the accent features corresponding to the multiple sub-speech information respectively, and the initial feature classification sub-model is configured to determine the to-be-determined accent type according to the accent features. The initial feature extraction sub-model includes a linear layer, and the linear layer is configured to perform linear classification on the accent features based on relevant parameters. The first adjustment unit is specifically configured to: Adjust the parameters of the initial feature extraction sub-model according to the first loss function and the second loss function to obtain a feature extraction sub-model, including: determining a comprehensive loss function according to the first loss function, the first weight parameter, the second loss function, and the second weight parameter; adjusting the parameters of the linear layer of the initial feature extraction sub-model according to the comprehensive loss function to obtain a feature extraction sub-model; Adjust the parameters of the initial feature classification sub-model according to the second loss function to obtain a feature classification sub-model; Determine the accent classification model according to the feature extraction sub-model and the feature classification sub-model.
7. The device according to claim 6, characterized in that The apparatus further includes a second acquisition unit, a fourth determination unit, a fifth determination unit, a sixth determination unit, and a second adjustment unit: The second acquisition unit is configured to acquire a self-supervised learning model and a second speech training sample; The fourth determination unit is configured to determine, through the self-supervised learning model, the target speech information corresponding to the second speech training sample, and the target speech information corresponds to the target speech part in the second speech training sample; The fifth determination unit is configured to determine, through the self-supervised learning model, the to-be-determined speech information corresponding to the target speech part according to the second speech training sample after removing the target speech information; The sixth determination unit is configured to determine the third loss function corresponding to the second speech training sample according to the difference between the target speech information and the to-be-determined speech information; The second adjustment unit is configured to adjust the parameters of the self-supervised learning model according to the third loss function and determine the initial feature extraction sub-model according to the adjusted self-supervised learning model.
8. The device according to claim 6, wherein The first adjustment unit is specifically configured to: In response to the number of times of parameter adjustment for the initial feature classification sub-model reaching a preset number of times, the parameters of the initial feature extraction sub-model are adjusted according to the first loss function and the second loss function to obtain the feature extraction sub-model.
9. A computer device, characterized in that, The device includes a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is used to execute the data processing method according to any one of claims 1-5 according to the instructions in the program code.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, and the computer program is used to execute the data processing method according to any one of claims 1-5.
11. A computer program product including instructions, when it runs on a computer, causes the computer to execute the data processing method according to any one of claims 1-5.