Method for generating training sample set and storage medium

CN115862607BActive Publication Date: 2026-09-04ALIBABA DAMO (HANGZHOU) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310099497.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-30
Publication Date
2026-09-04
Estimated Expiration
2043-01-30

AI Technical Summary

Technical Problem

[0005]本申请实施例提供了一种训练样本集的生成方法和存储介质,以至少解决无法有效处理训练样本集的技术问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115862607B_ABST
    Figure CN115862607B_ABST
Patent Text Reader

Abstract

The application discloses a method for generating a training sample set and a storage medium. The method comprises the following steps: obtaining an original training sample set to be processed, wherein the original training sample set is used for training a speech processing model; performing mixed enhancement on the original training sample to obtain a first target training sample set; performing contrast learning on the first target training sample set and the original training sample set to obtain a contrast loss, wherein the contrast loss is used for representing the similarity of the original training sample set relative to the first target training sample set; and adjusting the first target training sample set based on at least the contrast loss to obtain a second target training sample set, wherein the similarity of the original training sample set relative to the second target training sample set is greater than a similarity threshold, and the second target training sample set is used for training the speech processing model. The application solves the technical problem that the training sample set cannot be effectively processed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and more specifically, to a method for generating a training sample set and a storage medium. Background Technology

[0002] Currently, the voice command systems built into many smart devices have brought greater convenience to our lives, allowing us to activate devices effortlessly with a simple command. For example, we can activate a mobile phone by saying "Hey, Xiao Ai".

[0003] In related technologies, smart devices typically generate audio representations of speech data based on speech processing models (such as voice wake-up models). The performance of these speech processing models is positively correlated with the training data, and generally, speech processing models require a large amount of data for training. However, with the increasing demand for personalization in smart devices, speech processing models struggle to be effectively trained under limited resources, resulting in the technical problem of being unable to effectively process training sample sets.

[0004] There is currently no effective solution to the above problems. Summary of the Invention

[0005] This application provides a method for generating a training sample set and a storage medium to at least solve the technical problem of not being able to effectively process the training sample set.

[0006] According to one aspect of the embodiments of this application, a method for generating a training sample set is provided. The method may include: acquiring an original training sample set to be processed, wherein the original training sample set is used to train a speech processing model; performing hybrid enhancement on the original training sample set to obtain a first target training sample set; performing comparative learning on the first target training sample set and the original training sample set to obtain a comparative loss, wherein the comparative loss is used to characterize the similarity between the original training sample set and the first target training sample set; adjusting the first target training sample set at least based on the comparative loss to obtain a second target training sample set, wherein the similarity between the original training sample set and the second target training sample set is greater than a similarity threshold, and the second target training sample set is used to train the speech processing model.

[0007] According to another aspect of the embodiments of this application, another method for determining a speech processing model is provided. The method may include: acquiring a target training sample set, wherein the target training sample set is obtained by adjusting a first target training sample set at least based on a contrastive loss, the contrastive loss being obtained by comparative learning of the first target training sample set and an original training sample set, and used to characterize the similarity between the original training sample set and the first target training sample set, the first target training sample set being obtained by performing hybrid enhancement on the original training sample set; and training a speech processing model based on the target training sample set in response to a similarity threshold between the original training sample set and the target training sample set.

[0008] According to another aspect of the embodiments of this application, another speech processing method is provided. The method may include: acquiring speech to be detected sent to a client; extracting at least one target keyword from the speech to be detected using a speech processing model, wherein the speech processing model is trained based on a target training sample set, the target training sample set is obtained by adjusting a first target training sample set at least based on a contrastive loss, the contrastive loss is obtained by comparative learning of the first target training sample set and the original training sample set, and is used to characterize the similarity between the original training sample set and the first target training sample set, the first target training sample set is obtained by hybrid enhancement of the original training sample set, and the similarity between the original training sample set and the target training set is greater than a similarity threshold; and activating the client based on the target keyword.

[0009] According to another aspect of the embodiments of this application, another method for generating training samples is provided. The method may include: obtaining a raw training sample set to be processed by calling a first interface, wherein the raw training sample set is used to train a speech processing model, the first interface includes a first parameter, the parameter value of the first parameter being the raw training sample set; performing hybrid enhancement on the raw training sample set to obtain a first target training sample set; performing comparative learning on the first target training sample set and the raw training sample set to obtain a comparative loss, wherein the comparative loss is used to characterize the similarity between the raw training sample set and the first target training sample set; adjusting the first target training sample set at least based on the comparative loss to obtain a second target training sample set, wherein the similarity between the raw training sample set and the second target training sample set is greater than a similarity threshold; and outputting a second target training sample set by calling a second interface, wherein the output second target training sample set is used to train a speech processing model, the second interface includes a second parameter, the parameter value of the second parameter being the second target training sample set.

[0010] According to one aspect of the embodiments of this application, an apparatus for generating a training sample set is provided. The apparatus may include: a first acquisition unit for acquiring an original training sample set to be processed, wherein the original training sample set is used to train a speech processing model; a first processing unit for performing hybrid enhancement on the original training sample set to obtain a first target training sample set; a second processing unit for performing comparative learning on the first target training sample set and the original training sample set to obtain a contrastive loss, wherein the contrastive loss is used to characterize the similarity between the original training sample set and the first target training sample set; and a third processing unit for adjusting the first target training sample set at least based on the contrastive loss to obtain a second target training sample set, wherein the similarity between the original training sample set and the second target training sample set is greater than a similarity threshold, and the second target training sample set is used to train the speech processing model.

[0011] According to another aspect of the embodiments of this application, a different speech processing model determination apparatus is provided. The apparatus may include: a second acquisition unit, configured to acquire a target training sample set, wherein the target training sample set is obtained by adjusting a first target training sample set at least based on a contrastive loss, the contrastive loss being obtained by comparative learning of the first target training sample set and an original training sample set, and used to characterize the similarity between the original training sample set and the first target training sample set, the first target training sample set being obtained by performing hybrid enhancement on the original training sample set; and a training unit, configured to train a speech processing model based on the target training sample set in response to a similarity threshold between the original training sample set and the target training sample set.

[0012] According to another aspect of the embodiments of this application, another speech processing apparatus is provided. The apparatus may include: a collection unit for collecting speech to be detected sent to a client; an extraction unit for extracting at least one target keyword from the speech to be detected using a speech processing model, wherein the speech processing model is trained based on a target training sample set, the target training sample set is obtained by adjusting a first target training sample set at least based on a contrastive loss, the contrastive loss is obtained by comparative learning of the first target training sample set and the original training sample set, and is used to characterize the similarity between the original training sample set and the first target training sample set, the first target training sample set is obtained by hybrid enhancement of the original training sample set, and the similarity between the original training sample set and the target training set is greater than a similarity threshold; and an activation unit for activating the client based on the target keyword.

[0013] According to another aspect of the embodiments of this application, another training sample generation apparatus is provided. The apparatus may include: a third acquisition unit, configured to acquire a raw training sample set to be processed by calling a first interface, wherein the raw training sample set is used to train a speech processing model, the first interface includes a first parameter, the parameter value of the first parameter being the raw training sample set; a fourth processing unit, configured to perform hybrid enhancement on the raw training sample set to obtain a first target training sample set; a fifth processing unit, configured to perform comparative learning on the first target training sample set and the raw training sample set to obtain a comparative loss, wherein the comparative loss is used to characterize the similarity between the raw training sample set and the first target training sample set; an adjustment unit, configured to adjust the first target training sample set at least based on the comparative loss to obtain a second target training sample set, wherein the similarity between the raw training sample set and the second target training sample set is greater than a similarity threshold; and an output unit, configured to output a second target training sample set by calling a second interface, wherein the output second target training sample set is used to train a speech processing model, the second interface includes a second parameter, the parameter value of the second parameter being the second target training sample set.

[0014] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored program, wherein, when the program is running, the device where the storage medium is located executes the method for generating the training sample set of any of the above-mentioned methods.

[0015] According to another aspect of the embodiments of this application, a processor is also provided, which is used to run a program, wherein the method for generating a training sample set of any of the above-mentioned methods is executed during program execution.

[0016] In this embodiment, an original training sample set to be processed is obtained, wherein the original training sample set is used to train a speech processing model; the original training samples are hybridized and enhanced to obtain a first target training sample set; the first target training sample set and the original training sample set are compared and learned to obtain a contrastive loss, wherein the contrastive loss is used to characterize the similarity between the original training sample set and the first target training sample set; the first target training sample set is adjusted based at least on the contrastive loss to obtain a second target training sample set, wherein the similarity between the original training sample set and the second target training sample set is greater than a similarity threshold, and the second target training sample set is used to train the speech processing model. In other words, the embodiments of this application perform hybrid enhancement on the original training sample set to obtain a first target training sample set. By comparing and learning the first target training sample set and the original training sample set, an auxiliary contrast loss is introduced to maximize the similarity between the original training sample set (original pre-mixed sample set) and the enhanced sample set (first target training sample set). This can increase the speech processing model's cognition of feature content, achieving the goal of improving the performance of the speech processing model under low resource conditions of various sizes. In this way, the technical effect of effectively processing the training sample set is achieved, solving the technical problem of not being able to effectively process the training sample set.

[0017] It is worth noting that the general description above and the detailed description that follow are merely for illustrative purposes and do not constitute a limitation on this application. Attached Figure Description

[0018] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0019] Figure 1 This is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a method for generating a training sample set according to an embodiment of this application.

[0020] Figure 2 This is a structural block diagram of a computing environment according to an embodiment of this application;

[0021] Figure 3 This is a structural block diagram of a service mesh according to an embodiment of this application;

[0022] Figure 4 This is a flowchart of a method for generating a training sample set according to an embodiment of this application;

[0023] Figure 5 This is a flowchart of another method for generating a training sample set according to an embodiment of this application;

[0024] Figure 6 This is a flowchart of another method for generating a training sample set according to an embodiment of this application;

[0025] Figure 7 This is a flowchart of another method for generating a training sample set according to an embodiment of this application;

[0026] Figure 8 This is a schematic diagram illustrating a computer device accessing a private network according to an embodiment of this application;

[0027] Figure 9 This is a schematic diagram of a contrastive learning speech fusion model architecture according to an embodiment of this application;

[0028] Figure 10 This is a schematic diagram of a dimensionality reduction algorithm that compares embeddings of different technologies according to an embodiment of this application;

[0029] Figure 11 This is a schematic diagram of a method for generating a training sample set according to an embodiment of this application;

[0030] Figure 12 This is a schematic diagram of a speech processing model determination device according to an embodiment of this application;

[0031] Figure 13 This is a schematic diagram of a voice processing device according to an embodiment of this application;

[0032] Figure 14 This is a schematic diagram of another training sample generation apparatus according to an embodiment of this application;

[0033] Figure 15 This is a structural block diagram of a computer terminal according to an embodiment of this application. Detailed Implementation

[0034] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0035] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0036] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows:

[0037] Keyword Spotting (KWS) refers to the real-time detection of a specific segment of a speaker in a continuous speech stream. It can be used to activate a device from a sleep state to a running state.

[0038] Data augmentation is a technique that uses algorithms to expand training data, or it can be used to automatically augment training data using algorithms.

[0039] Contrastive learning, a type of self-supervised learning, refers to automatically constructing similar and dissimilar instances and learning a representation model. Through this model, similar instances can be made closer in the projection space, while dissimilar instances can be made farther apart.

[0040] Overfitting can refer to the phenomenon where a particular dataset is matched too closely or precisely to other data or to predict future observations.

[0041] Regularization can refer to the process of adding extra information to solve adaptive problems or overfitting. In the optimization process of machine learning and inverse problems, regularization can be added to the objective function.

[0042] Example 1

[0043] According to an embodiment of this application, a method for generating a training sample set is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0044] The method embodiment provided in Embodiment 1 of this application can be executed on a mobile terminal, computer terminal, or similar computing device. Figure 1 This is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a method for generating a training sample set, according to an embodiment of this application. Figure 1 As shown, computer terminal A (or mobile device) may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a Universal Serial Bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal A may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0045] It should be noted that the aforementioned processor 102 and / or other training sample set generation circuitry are generally referred to herein as "training sample set generation circuitry". This training sample set generation circuitry can be wholly or partially embodied in software, hardware, firmware, or any other combination thereof. Furthermore, the training sample set generation circuitry can be a single, independent processing module, or wholly or partially integrated into any other element in computer terminal A (or mobile device). As involved in the embodiments of this application, the training sample set generation circuitry serves as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).

[0046] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the training sample set generation method in the embodiments of this application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the above-mentioned training sample set generation method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to computer terminal A via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0047] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of computer terminal A. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.

[0048] The display can be, for example, a touchscreen liquid crystal display (LCD), which allows the user to interact with the user interface of a computer terminal A (or mobile device).

[0049] Figure 1 The hardware structure block diagram shown can serve not only as an exemplary block diagram of the aforementioned computer terminal A (or mobile device), but also as an exemplary block diagram of the aforementioned server. In one optional embodiment, Figure 2 The use of the above is illustrated in a block diagram. Figure 1 The computer terminal A (or mobile device) shown is an example of a computing node in computing environment 201. Figure 2 This is a structural block diagram of a computing environment according to an embodiment of this application, such as... Figure 2As shown, computing environment 201 includes multiple computing nodes (such as servers) running on a distributed network (shown as 210-1, 210-2, ..., in the diagram). Each computing node contains local processing and memory resources, and end user 202 can remotely run applications or store data within computing environment 201. Applications can be provided as multiple services 220-1, 220-2, 220-3, and 220-4 within computing environment 201, representing services "A", "D", "E", and "H", respectively.

[0050] End user 202 can provide and access services through a web browser or other software application on a client. In some embodiments, the provisioning and / or requests of end user 202 can be provided to ingress gateway 230. Ingress gateway 230 may include a corresponding agent to handle the provisioning and / or requests for services (one or more services provided in computing environment 201).

[0051] Services are provided or deployed based on various virtualization technologies supported by the computing environment 201. In some embodiments, services may be provided based on virtual machine (VM)-based virtualization, container-based virtualization, and / or similar methods. VM-based virtualization can simulate a real computer by initializing a virtual machine, executing programs and applications without directly accessing any actual hardware resources. While the machine is virtualized by a virtual machine, container-based virtualization can launch containers to virtualize an entire operating system (OS), allowing multiple workloads to run on a single OS instance.

[0052] In one embodiment based on container virtualization, several containers of a service can be assembled into a Pod (e.g., a Kubernetes Pod). For example, such as Figure 2 As shown, service 220-2 can be equipped with one or more Pods 240-1, 240-2, ..., 240-N (collectively referred to as Pods). Each Pod can include a proxy 245 and one or more containers 242-1, 242-2, ..., 242-M (collectively referred to as containers). One or more containers in a Pod handle requests related to one or more corresponding functions of the service. The proxy 245 typically controls service-related network functions such as routing and load balancing. Other services can also be equipped with Pods similar to Pods.

[0053] During operation, executing a user request from end user 202 may require calling one or more services in computing environment 201. Executing one or more functions of one service may require calling one or more functions of another service. For example... Figure 2As shown, service "A" 220-1 receives user requests from terminal user 202 from ingress gateway 230. Service "A" 220-1 can call service "D" 220-2, and service "D" 220-2 can request service "E" 220-3 to perform one or more functions.

[0054] The aforementioned computing environment can be a cloud computing environment, where resource allocation is managed by cloud services, allowing functionality development without needing to consider implementation, adjustment, or server scaling. This computing environment allows developers to execute event-responsive code without building or maintaining complex infrastructure. Services can be partitioned into a set of functions that can automatically and independently scale, rather than scaling a single hardware device to handle potential loads.

[0055] In another alternative embodiment, Figure 3 The use of the above is illustrated in a block diagram. Figure 1 The computer terminal A (or mobile device) shown is an example of a service mesh. Figure 3 This is a structural block diagram of a service mesh according to an embodiment of this application, such as... Figure 3 As shown, the service mesh 300 is mainly used to facilitate secure and reliable communication between multiple microservices. Microservices refer to the decomposition of an application into multiple smaller services or instances, which are distributed across different clusters / machines.

[0056] like Figure 3 As shown, a microservice may include application service instance A and application service instance B, which together form the functional application layer of service mesh 300. In one implementation, application service instance A runs as a container / process 308 on machine / workload container group 314 (Pod), and application service instance B runs as a container / process 310 on machine / workload container group 316 (Pod).

[0057] In one implementation, application service instance A can be a product query service, and application service instance B can be a product order placement service.

[0058] like Figure 3As shown, application service instance A and grid agent (sidecar) 303 coexist in machine workload container group 314, and application service instance B and grid agent 305 coexist in machine workload container 314. Grid agent 303 and grid agent 305 form the data plane layer of service mesh 300. Grid agent 303 and grid agent 305 run as container / process 304 and container / process 306 respectively, and can receive requests 312 for product query services. Grid agent 303 and application service instance A can communicate bidirectionally, and grid agent 305 and application service instance B can also communicate bidirectionally. Furthermore, grid agent 303 and grid agent 305 can also communicate bidirectionally with each other.

[0059] In one implementation, all traffic from application service instance A is routed to the appropriate destination via mesh proxy 303, and all network traffic from application service instance B is routed to the appropriate destination via mesh proxy 305. It should be noted that the network traffic mentioned herein includes, but is not limited to, Hypertext Transfer Protocol (HTTP), Representational State Transfer (REST), and other high-performance protocols.

[0060] In one implementation, the functionality of the extended data plane layer can be achieved by writing custom filters for the proxy (Envoy) in service mesh 300. The service mesh proxy configuration can enable the service mesh to correctly proxy service traffic, achieving service interoperability and service governance. Mesh proxy 303 and mesh proxy 305 can be configured to perform at least one of the following functions: service discovery, health checking, routing, load balancing, authentication and authorization, and observability.

[0061] like Figure 3 As shown, the service mesh 300 also includes a control plane layer. This control plane layer can consist of a set of services running in a dedicated namespace, hosted by a managed control plane component 301 within machine / workload container groups (machine / Pods) 302. Figure 3As shown, the managed control plane component 301 communicates bidirectionally with grid agents 303 and 305. The managed control plane component 301 is configured to perform various control and management functions. For example, it receives telemetry data from grid agents 303 and 305 and can further aggregate this telemetry data. In addition to these services, the managed control plane component 301 can also provide a user-facing Application Programming Interface (API) to facilitate manipulation of network behavior and the provision of configuration data to grid agents 303 and 305.

[0062] Under the aforementioned operating environment, this application provides the following: Figure 4 The method for generating the training sample set is shown. Figure 4 This is a flowchart of a method for generating a training sample set according to an embodiment of this application. Figure 4 As shown, the method may include the following steps:

[0063] Step S402: Obtain the original training sample set to be processed, wherein the original training sample set is used to train the speech processing model.

[0064] In the technical solution provided in step S402 of this application, an original training sample set to be processed can be obtained. This original training sample set may include a sample set of speech data, such as a very small training set (e.g., 2.5 mins, 5 mins, 10 mins) or a small number of user samples, or it may be a pre-mixed set (pre-mixed sample) obtained by mixing, or it may be audio, video, etc. containing speech data, denoted by Xi and Xj. This is only an example and does not impose specific limitations on the content and representation of the original training sample set. The speech processing model can be a Keyword Spotting (KWS) model, which can be used to process speech data.

[0065] Optionally, after obtaining the original training sample set to be processed, a speech processing model can be trained using the original training sample set.

[0066] For example, a voice command dataset from a client, such as a browser, can be obtained. This dataset contains a total of 105,000 utterances, including 35 unique words. Each audio sample in the dataset is stored as a one-second or shorter audio file (e.g., WAV format) sampled at 16kHz, thus obtaining the raw training sample set to be processed.

[0067] Step S404: Perform hybrid enhancement on the original training sample set to obtain the first target training sample set.

[0068] In the technical solution provided by step S404 of this application, the original training samples can be mixed-up augmented to obtain a first target training sample set. The mixed-up augmentation can be used to generate content-rich speech codes. (First target training sample set) It can be used as a training sample or as an augmented sample set.

[0069] Optionally, the samples in the original training sample set can be enhanced by mixing small noise distortion, time shift, time stretching and spectral enhancement to obtain the first target training sample set. It should be noted that the mixed enhancement method here is only an example and no specific limitation is made on the mixed enhancement method.

[0070] For example, a random time offset of -100 milliseconds to 100 milliseconds can be applied to a portion of the original training samples in the original training sample set, and a random time stretch of 0.9 to 1.1 coefficients can be applied to a randomly selected portion of the original training samples. Spectral enhancement with a time masking size of 13 and a spectral masking size of 7 can be applied to the original training samples to obtain the first target training sample set.

[0071] Step S406: Perform comparative learning on the first target training sample set and the original training sample set to obtain the comparative loss, wherein the comparative loss is used to characterize the similarity between the original training sample set and the first target training sample set.

[0072] In the technical solution provided by step S406 of this application, comparative learning can be performed on the first target training sample set and the original training sample set to obtain a comparative loss. The comparative loss can be used to characterize the similarity (relative similarity) between the original training sample set and the first wooden plaque training sample set, and can serve as an auxiliary comparative loss.

[0073] Because the training samples in the first target training sample set obtained after hybrid enhancement may generate highly distorted signals from two overlapping speech samples—for example, there may be noisy data (specific spikes in model error)—these can dominate the effective gradient and affect the training of the speech processing model. To address this issue, in this embodiment, comparative learning is performed on the first target training sample set and the original training sample set. By using comparative learning, the relative similarity between the samples in the first target training sample set and the samples in the original training sample set is maximized, thereby improving the accuracy of speech processing model training.

[0074] For example, an auxiliary component can be used to perform comparative learning on the first target training sample set and the original training sample set to obtain a comparative loss. The comparative learning can then maximize the similarity between the first target training sample set and the original training sample set.

[0075] Step S408: Adjust the first target training sample set based at least on the contrast loss to obtain the second target training sample set, wherein the similarity between the original training sample set and the second target training sample set is greater than the similarity threshold, and the second target training sample set is used to train the speech processing model.

[0076] In the technical solution provided in step S408 of this application, the first target training sample set can be adjusted based at least on the contrastive loss to obtain the second target training sample set. The similarity between the original training sample set and the second target training sample set is greater than a similarity threshold. The second target training sample set can be used to train a speech processing model. The similarity threshold can be a preset similarity threshold, for example, 50%. This is merely an example, and no specific limitation is made on the size of the similarity threshold.

[0077] In this embodiment, the first target sample set can be adjusted based on the contrast loss to obtain a second target training sample set that is closer to the original training sample set. The second target training sample set can be used as training data to train the speech processing model.

[0078] In related technologies, with the increasing demand for personalized smart devices, customized speech processing models are needed to quickly adapt to limited original training sample sets (user samples). However, when the number of user samples is limited, the speech processing model is easily undertrained, resulting in low prediction accuracy. In this application, the embodiment obtains a first target training sample set by performing hybrid enhancement on the original training sample set, and then adjusts the first target training sample set to obtain a second target training sample set. Based on the original training sample set, the first target training sample set, and the second target training sample set, the speech processing model is trained using these three training sample sets. This addresses the issue of the speech processing model's generalization ability under low-resource conditions (e.g., limited user samples, small training data). Furthermore, the enhancement technique injects slight variability into the original training sample set (data instances) to better prevent the speech processing model from memorizing the dataset, thereby reducing overfitting.

[0079] Through steps S402 to S408 of this application, an original training sample set to be processed is obtained, wherein the original training sample set is used to train a speech processing model; the original training samples are hybridized and enhanced to obtain a first target training sample set; the first target training sample set and the original training sample set are compared and learned to obtain a contrastive loss, wherein the contrastive loss is used to characterize the similarity between the original training sample set and the first target training sample set; the first target training sample set is adjusted based at least on the contrastive loss to obtain a second target training sample set, wherein the similarity between the original training sample set and the second target training sample set is greater than a similarity threshold, and the second target training sample set is used to train a speech processing model. In other words, the embodiments of this application perform hybrid enhancement on the original training sample set to obtain a first target training sample set. By comparing and learning the first target training sample set and the original training sample set, an auxiliary contrast loss is introduced to maximize the similarity between the original training sample set (original pre-mixed sample set) and the enhanced sample set (first target training sample set). This can increase the speech processing model's cognition of feature content, achieving the goal of improving the performance of the speech processing model under low resource conditions of various sizes. In this way, the technical effect of effectively processing the training sample set is achieved, solving the technical problem of not being able to effectively process the training sample set.

[0080] The method described in this embodiment will be further described below.

[0081] As an optional implementation, step S406 involves comparative learning of the first target training sample set and the original training sample set to obtain a comparative loss, including: mapping a vector composed of speech data from the original training sample set to a first target vector of the target dimension; mapping a vector composed of speech data from the first target training sample set to a second target vector of the target dimension, wherein the speech data in the first target training sample set is obtained by mixing and enhancing the speech data from the original training sample set; and comparative learning of the first target vector and the second target vector to obtain a comparative loss.

[0082] In this embodiment, the vectors formed by the speech data in the original training sample set can be mapped to obtain a first target vector of the target dimension. Furthermore, the speech data in the original training sample set can be hybridized and enhanced to obtain a first target training sample in the first target training sample set. The vectors formed by the speech data in the first target training sample can then be mapped to obtain a second target vector of the target dimension. A contrastive learning process can be performed between the first target vector and the second target vector to obtain a contrastive loss. The target dimension can be a dimension selected according to the actual situation, for example, it can be 128. The first target vector (f) p (X rThis can be used to characterize the content of speech data in the original training sample set. Second target vector It can be used to characterize the content of speech data in the first target training sample.

[0083] Optionally, an original training sample set can be obtained, and the speech data in the original training sample set can be hybridized and enhanced to obtain a first target training sample set. Both the original training sample set and the first target training sample set can be sample sets containing at least one speech data point, such as multiple audio segments. This is merely an example, and no specific restrictions are placed on the content of the samples in the original training sample set and the first target training sample set. The vectors (latent vectors) formed by the speech data in the original training sample set and the first target training sample set can be mapped separately to obtain a first target vector and a second target vector in the target dimension. The first target vector and the second target vector can then be compared and learned to obtain a contrastive loss.

[0084] For example, two samples (Xi and Xj) can be arbitrarily selected from the original training sample set, and the two selected samples can be mixed and enhanced to obtain the training samples in the first target training sample set. Xi, Xj, and The data is passed to a shared encoding network to obtain vector representations of the speech data from the three training samples. These vectors can then be projected to obtain a first target vector and a second target vector in the target dimension. A contrastive loss can then be obtained by comparing and contrasting the first and second target vectors.

[0085] As an optional implementation, comparative learning is performed on the first target vector and the second target vector to obtain the comparative loss, including: obtaining the norm value of the first target vector and the norm value of the second target vector, wherein the norm value of the first target vector is used to represent the length of the first target vector and the norm value of the second target vector is used to represent the length of the second target vector; and determining the comparative loss based on the norm value of the first target vector and the norm value of the second target vector.

[0086] In this embodiment, the norm values ​​of the first target vector and the second target vector can be obtained, and the contrast loss can be determined based on the norm values ​​of the first and second target vectors. Wherein, the norm value of the first target vector (||f) p (X r ))||2) can be used to represent the length of the first target vector and the norm of the second target vector. It can be used to represent the length of the second target vector.

[0087] Optionally, in order to calculate the loss of contrastive learning, the obtained first target vector and second target vector can be subjected to L2 norm to obtain the norm value of the first target vector and the norm value of the second target vector. The contrastive loss can be determined based on the norm value of the first target vector and the norm value of the second target vector.

[0088] As an optional implementation, determining the contrast loss based on the norm of the first target vector and the norm of the second target vector includes: obtaining the mean square error between the second target vector and the first target vector, wherein the mean square error is used to represent the degree of difference between the first target vector and the second target vector; and determining the contrast loss based on the mean square error, the norm of the first target vector, and the norm of the second target vector.

[0089] In this embodiment, the mean squared error between the first target vector and the second target vector can be obtained. The contrast loss can be determined based on the mean squared error, the norm of the first target vector, and the norm of the second target vector. The mean squared error is... It can be used to represent the degree of difference between the first target vector and the second target vector.

[0090] Optionally, to determine the contrast loss, the L2 norm can be applied to the first and second target vectors. The similarity between the normalized first and second target vectors can be calculated using the mean squared error. The contrast loss can then be determined using the following formula.

[0091]

[0092] Here, fp(.) is the projection of the speech processing model onto r∈{i,j}.

[0093] As an optional implementation, mapping a vector composed of speech data from the original training sample set to a first target vector of the target dimension includes: in an encoding network model, mapping a vector composed of speech data from the original training sample set to a first target vector of the target dimension, wherein the encoding network model is used to encode the input vector into a vector of the target dimension; mapping a vector composed of speech data from the first target training sample to a second target vector of the target dimension includes: in an encoding network model, mapping a vector composed of speech data from the first target training sample to a second target vector of the target dimension.

[0094] In this embodiment, the original training sample set and the first target training sample set can be input into the encoding network model. In the encoding network model, the vector composed of speech data from the original training sample set is mapped to a first target vector of the target dimension, and the vector composed of speech data from the first target training sample set is mapped to a second target vector of the target dimension, thus obtaining the first target vector and the second target vector. The encoding network model can be a shared encoding network, which can be used to encode the input vector into a vector of the target dimension.

[0095] Optionally, the speech data from the original training sample set and the speech data from the first target training sample set can be input into the encoding network model, and the vector formed by the speech data can be mapped in the encoding network model to obtain the first target vector and the second target vector.

[0096] As an optional implementation, the cross-entropy loss of the speech data in the original training samples is determined; the first target training sample set is adjusted based at least on the contrastive loss to obtain the second target training sample set, including: adjusting the first target training sample set based on the cross-entropy loss and the contrastive loss to obtain the second target training sample set.

[0097] In this embodiment, the cross-entropy loss of the speech data in the original training samples is determined. Based on the cross-entropy loss and the contrast loss, the first target training samples can be adjusted to obtain the second target training sample set.

[0098] For example, hybrid augmentation can employ the principle of minimizing neighbor risk to encourage the classifier to perform linearly on the training samples. This principle of minimizing neighbor risk reduces unwanted variability when performing inference on unseen speech data. During data preprocessing, two speech data points (audio instances) can be randomly selected from the original training sample set, and a dummy training example can be constructed using the following formula:

[0099]

[0100]

[0101] Where i, j∈1; N can be used to represent the index of the original training sample set; x and y can represent the original input waveform and one-hot label encoding, respectively; λ~Beta(α, α), for α∈(0,∞) can be the interpolation parameter to determine the amount of content to be linearly mixed, given the virtual input label pair It can be a sample from the training set of the first target sample after hybrid enhancement. The time-spectral characteristics of the Short Time Fourier Transform (STFT) can be used to calculate the cross-entropy (L). mix The losses are as follows:

[0102]

[0103] Where f(.) can be the acoustic coding model, and CE can refer to the standard cross-entropy loss.

[0104] Optionally, the first target training sample set can be adjusted based on the calculated cross-entropy loss and contrastive loss to obtain the second target training sample set.

[0105] As an optional implementation, the first target training sample set is adjusted based on cross-entropy loss and contrastive loss to obtain the second target training sample set, including: adjusting the contrastive loss to the target loss based on cross-entropy loss; and adjusting the first target training sample set based on the target loss to obtain the second target training sample set.

[0106] In this embodiment, the contrastive loss can be adjusted based on the cross-entropy loss to achieve the purpose of adjusting the contrastive loss to the target loss. The first target training sample set can be adjusted based on the target loss to obtain the second target training sample set.

[0107] Alternatively, the contrast loss can be adjusted by addition, and the target loss (L) can be determined by the following formula:

[0108]

[0109] Here, β can be a penalty parameter that measures the contribution of the contrastive loss. It can be determined through testing; for example, it can be set to 0.5. This is just an example, and no specific limit is placed on the size of the penalty parameter. Λ r This can be used to assign weights to the comparison loss.

[0110] Alternatively, the weights of the contrastive loss can be calculated using the following formula:

[0111]

[0112] As an optional implementation, random data augmentation is performed on the speech data in the original training samples; comparative learning is performed on the first target training sample set and the original training sample set to obtain a comparative loss, including: comparative learning is performed on the first target training sample set and the augmented original training sample set to obtain a comparative loss.

[0113] In this embodiment, random data augmentation can be performed on the speech data in the original training samples, and comparative learning can be performed on the augmented first target training sample set and the original training sample set to obtain a contrastive loss.

[0114] This application also proposes a method for audio fusion enhancement and contrastive learning for mixed audio data. This method can enhance content-rich speech coding by using consistent attributes between positive enhancement samples. Furthermore, by using a first target training sample after random data enhancement to train the model, an embedding with lower complexity and higher guarantee can be obtained. This enables the trained speech processing model to generate more effective representations and achieve greater generalization ability with low resources. This solves the technical problem of not being able to effectively process the training sample set and improves the training efficiency of the model.

[0115] Optionally, by comparing the first target training sample set with the original training samples, a relative contrastive loss can be introduced, which can reduce the variability of perturbation samples, achieve less complex embedding, and thus improve the training efficiency of the speech processing model under low resource conditions.

[0116] As an optional implementation, data augmentation is performed on the speech data in the original training sample set, including: random data augmentation of the speech data in the original training samples.

[0117] In this embodiment, random data augmentation can be performed on the speech data in the original training samples. Random data augmentation can involve randomly selecting data augmentation methods from the speech data. These methods can include time shifting, time stretching, etc.; here, only distance is considered, and no specific limitations are imposed on the data augmentation methods.

[0118] For example, two speech data points can be randomly selected from the original training sample set, or original training samples that have undergone data augmentation such as time shifting or time stretching can be randomly selected from the original training sample set. The method of data augmentation for the original speech samples can be randomly chosen, thereby enhancing the flexibility of the model's detection.

[0119] As an optional implementation, the original training sample set is mixed and enhanced to obtain a first target training sample set, including: mixing the speech data in the original training samples to obtain mixed speech data; and performing random data augmentation on the mixed speech data to obtain the first target training sample set.

[0120] In this embodiment, the speech data in the original training samples can be mixed to obtain mixed speech data, and random data augmentation can be performed on the mixed speech data to obtain the first target training sample set.

[0121] Optionally, the first target training sample includes speech augmentation data obtained by randomly augmenting a single speech data and speech augmentation data obtained by randomly augmenting multiple mixed speech data.

[0122] For example, random data augmentation can be performed on speech data A and speech data B in the original training samples, and speech data A and speech data B can be mixed to obtain mixed speech data AB. Random data augmentation can then be performed on the mixed speech data AB to obtain the first target training sample set. The first target training sample set includes samples after random data augmentation of speech data A, samples after random augmentation of speech data B, and samples after random augmentation of mixed speech data AB.

[0123] As an optional implementation, the original training sample set includes labeled data with a data volume less than a data volume threshold.

[0124] In this embodiment, the original training sample set may include labeled data with a data volume less than a data volume threshold. The data volume threshold can be a pre-set value, such as the data volume corresponding to an average of 2.5 minutes per word or an average of 5 minutes per word. This is merely an example, and no specific limitation is made on the size of the data volume threshold. The annotations in the labeled data can be based on the speaker, or on keywords, etc.

[0125] In this embodiment of the invention, by using different labeled data to train the model, not only is the diversity of the training process increased, but the learning difficulty is also increased, thereby improving the breadth of model deployment and the accuracy of prediction results.

[0126] In this embodiment, the original training sample set is hybridized and enhanced to obtain a first target training sample set. By comparing and learning the first target training sample set and the original training sample set, an auxiliary contrastive loss is introduced to maximize the similarity between the original training sample set (original premixed sample set) and the enhanced sample set (first target training sample set). This can increase the speech processing model's recognition of feature content and achieve the goal of improving the performance of the speech processing model under low resource conditions of various sizes. In this way, the technical effect of effectively processing the training sample set is achieved, solving the technical problem of not being able to effectively process the training sample set.

[0127] This application also provides another method for generating training sample sets, which can be applied to the model training side. Figure 5 This is a flowchart of another method for generating a training sample set according to an embodiment of this application, such as... Figure 5 As shown, the method may include the following steps.

[0128] Step S502: Obtain the target training sample set, wherein the target training sample set is obtained by adjusting the first target training sample set based at least on the contrast loss, the contrast loss is obtained by comparing and learning the first target training sample set and the original training sample set, and is used to characterize the similarity between the original training sample set and the first target training sample set, and the first target training sample set is obtained by performing hybrid enhancement on the original training samples.

[0129] In the technical solution provided in step S502 of this application, the original training samples are obtained, and the original training samples can be hybridized and enhanced to obtain a first target training sample set. Comparative learning can be performed on the first target training sample set and the original training sample set to obtain a contrastive loss used to characterize the similarity between the first target training sample set and the original training sample set. The first target training sample set can be adjusted based on the contrastive loss to obtain the target training sample set.

[0130] Step S504: In response to the similarity between the original training sample set and the target training sample set being greater than the similarity threshold, a speech processing model is trained based on the target training sample set.

[0131] In the technical solution provided in step S504 of this application, it is determined whether the similarity between the original training sample set and the target training sample set is greater than the similarity threshold. If the similarity between the original training sample set and the target training sample set is greater than the similarity threshold, a speech processing model can be trained based on the target training sample set.

[0132] Through steps S502 to S504 of this application, original training samples are obtained. These samples can be hybridized and enhanced to obtain a first target training sample set. Comparative learning can be performed on the first target training sample set and the original training sample set to obtain a contrastive loss representing the similarity between the two sets. The first target training sample set can be adjusted based on this contrastive loss to obtain a target training sample set. When the similarity between the original training sample set and the target training sample set exceeds a similarity threshold, a speech processing model can be trained based on the target training sample set. This achieves the technical effect of effectively processing the training sample set and solves the technical problem of ineffective training sample set processing.

[0133] This application also provides another method for generating training sample sets, which can be applied to scenarios such as speech recognition and voice wake-up. Figure 6 This is a flowchart of another method for generating a training sample set according to an embodiment of this application, such as... Figure 6 As shown, the method may include the following steps.

[0134] Step S602: Collect the voice to be detected sent to the client.

[0135] In the technical solution provided in step S602 of this application, the voice to be detected can be collected and sent to the client. The client can be a mobile client, such as a mobile phone or computer. The voice to be detected can be voice data emitted by a human voice, voice data emitted by a video, etc.; this is merely an example and does not impose specific limitations on the content of the voice to be detected.

[0136] For example, the audio to be detected can be obtained by capturing video sent to the client using a microphone or other audio capture device.

[0137] Step S604: Extract at least one target keyword from the speech to be detected using a speech processing model. The speech processing model is trained based on a target training sample set. The target training sample set is obtained by adjusting a first target training sample set based on at least a contrast loss. The contrast loss is obtained by comparing and learning the first target training sample set and the original training sample set, and is used to characterize the similarity between the original training sample set and the first target training sample set. The first target training sample set is obtained by performing hybrid enhancement on the original training samples. The similarity between the original training sample set and the target training sample set is greater than a similarity threshold.

[0138] In the technical solution provided in step S604 of this application, at least one target keyword can be extracted from the speech to be detected using a speech processing model. The target keyword can be a wake-up word, such as "Xiao Ai, open program A for me."

[0139] Optionally, the original training samples can be hybridized to obtain a first target training sample set. Comparative learning is then performed between the first target training sample set and the original training sample set to obtain a contrastive loss representing the original training sample set relative to the first target training sample set. The first target training sample set is then adjusted based on this contrastive loss to obtain a target training sample set. A speech processing model is then trained based on the target training sample set. This speech processing model processes the speech to be detected, extracting at least one target keyword. The target keyword can be the same as preset keywords, such as "Xiao Ai Tongxue" or "Xiao Du Tongxue," etc. This is merely an example and does not impose specific restrictions on the content of the keywords.

[0140] Step S606: Activate the client based on the target keyword.

[0141] In the technical solution provided by step S606 of this application, when a pre-set target keyword is detected in the speech to be detected, the client can be activated based on the target keyword.

[0142] For example, when the client is a mobile phone, when taking a photo, the target keyword can be set to "eggplant". The system can obtain the speech to be detected in the current environment. When "eggplant" is detected in the speech, the mobile phone can be activated to take a photo.

[0143] As an optional implementation, activating the client based on the target keyword includes: activating the client in response to the similarity between the target keyword and a predetermined keyword associated with the client being greater than a predetermined similarity threshold.

[0144] In this embodiment, the similarity between the target keyword and the pre-defined keywords associated with the client is determined. If the similarity between the target keyword and the pre-defined keywords associated with the client exceeds a pre-defined similarity threshold, the client can be activated. The pre-defined keywords can be pre-set keywords.

[0145] In this embodiment, the process involves collecting the speech to be detected sent to the client; extracting at least one target keyword from the speech using a speech processing model, wherein the speech processing model is trained based on a target training sample set, the target training sample set is obtained by adjusting a first target training sample set based at least on a contrastive loss, the contrastive loss is obtained by comparing the first target training sample set and the original training sample set, and is used to characterize the similarity between the original training sample set and the first target training sample set, the first target training sample set is obtained by performing hybrid enhancement on the original training samples, and the similarity between the original training sample set and the target training sample set is greater than a similarity threshold; and activating the client based on the target keyword, thereby achieving the technical effect of effectively processing the training sample set and solving the technical problem of not being able to effectively process the training sample set.

[0146] This application also provides another method for generating training sample sets, which can be applied to software-as-a-Service (SaaS). Figure 7 This is a flowchart of another method for generating a training sample set according to an embodiment of this application, such as... Figure 7 As shown, the method may include the following steps.

[0147] Step S702: Obtain the original training sample set to be processed by calling the first interface. The original training sample set is used to train the speech processing model. The first interface includes a first parameter, and the parameter value of the first parameter is the original training sample set.

[0148] In the technical solution provided in step S702 of this application, the first interface can be an interface for data interaction between the server and the user. The user can obtain the original training sample set to be processed by calling the first interface. The original training sample set is used as a first parameter of the first interface to achieve the purpose of obtaining the original training sample set to be processed.

[0149] Step S704: Perform hybrid enhancement on the original training samples to obtain the first target training sample set.

[0150] Step S706: Perform comparative learning on the first target training sample set and the original training sample set to obtain the comparative loss, wherein the comparative loss is used to characterize the similarity between the original training sample set and the first target training sample set.

[0151] Step S708: Adjust the first target training sample set based at least on the contrast loss to obtain the second target training sample set, wherein the similarity between the original training sample set and the second target training sample set is greater than the similarity threshold.

[0152] Step S710: Output the second target training sample set by calling the second interface. The output second target training sample set is used to train the speech processing model. The second interface includes a second parameter, and the parameter value of the second parameter is the second target training sample set.

[0153] In the technical solution provided by step S710 of this application, the second interface can be an interface for data interaction between the server and the user. The server can send the second target training sample set to the client, so that the client can output the second target training sample set to the second interface as a parameter of the second interface, thereby achieving the purpose of sending the second target training sample set to the user.

[0154] Figure 8 This is a schematic diagram illustrating a computer device accessing a private network according to an embodiment of this application, such as... Figure 8 As shown, the original training sample set to be processed can be obtained by calling the first interface. The computer device executes the following steps: Step S802, obtaining the original training sample set to be processed by calling the first interface; Step S804, performing hybrid enhancement on the original training sample set to obtain the first target training sample set; Step S806, performing comparative learning on the first target training sample set and the original training sample set to obtain the comparative loss; Step S808, adjusting the first target training sample set based at least on the comparative loss to obtain the second target training sample set; Step S810, outputting the second target training sample set by calling the second interface.

[0155] Through steps S702 to S708 of this application, the original training sample set to be processed is obtained by calling a first interface, wherein the original training sample set is used to train a speech processing model, and the first interface includes a first parameter whose parameter value is the original training sample set; the original training samples are hybridized and enhanced to obtain a first target training sample set; the first target training sample set and the original training sample set are compared and learned to obtain a contrastive loss, wherein the contrastive loss is used to characterize the similarity between the original training sample set and the first target training sample set; the first target training sample set is adjusted based at least on the contrastive loss to obtain a second target training sample set, wherein the similarity between the original training sample set and the second target training sample set is greater than a similarity threshold; the second target training sample set is output by calling a second interface, wherein the output second target training sample set is used to train a speech processing model, and the second interface includes a second parameter whose parameter value is the second target training sample set, thereby achieving the technical effect of effectively processing the training sample set and solving the technical problem of not being able to effectively process the training sample set.

[0156] Example 2

[0157] The voice command systems built into many smart devices have brought greater convenience to our lives, allowing us to easily activate devices with a simple wake-up command. For example, we can activate our mobile devices by saying "Hey, Xiao Ai".

[0158] Currently, systems capable of recognizing these commands require the establishment of a voice wake-up model to detect predetermined words in continuous speech. Operationally, this system can convert the raw audio into time-spectral features, then pass these features to an acoustic neural network to predict the keyword category that satisfies predetermined conditions and minimizes the error rate.

[0159] In related technologies, high-precision voice wake-up networks are typically created using deep acoustic neural networks (DNNs) by employing transformer blocks, convolutional blocks, autoregressive layers, or hybrid structures of these architectures. These neural network-based models usually require a large number of training samples to learn appropriate audio representations of predefined keywords, thus avoiding overfitting, especially for deeper networks. However, with the increasing demand for personalized smart devices, customized voice wake-up systems are needed to quickly adapt to limited user samples. But when the number of user samples is limited, the model is prone to inadequate training, leading to technical problems such as low prediction accuracy.

[0160] To address the aforementioned issues, various data augmentation techniques have been proposed, such as hybrid small noise distortion, time shifting, time stretching, and spectral enhancement. These techniques aim to improve the generalization ability of deep neural networks under low-resource conditions (e.g., limited user samples, small datasets). These augmentation techniques inject small variability into data instances and prevent the model from memorizing the dataset, thereby reducing overfitting. However, the limited types of perturbations available in speech preprocessing restrict the diversity of data augmentation.

[0161] To address the aforementioned issues, this application proposes a low-resource contrastive learning speech mixing (Cos-Mix) learning algorithm based on an input mixing algorithm. This algorithm introduces an auxiliary contrastive loss into existing mixing enhancement techniques to maximize the relative similarity between the original pre-mixed samples and the enhanced samples, thereby injecting enhancement constraints and guiding the model towards simpler yet richer content-based speech representations from the two enhanced views. The two enhanced views can be noisy mixed speech and clean pre-mixed speech. Through this method, an algorithm with good performance is trained on very small training sets (e.g., 2.5 minutes, 5 minutes, or 10 minutes), thus improving the accuracy of model training in low-resource speech application scenarios without increasing labeled data.

[0162] In low-resource contrastive learning algorithms for speech mixing, data augmentation (mix-up) is a regularization technique that fits a network of two linearly interpolated samples to their corresponding soft labels. However, the mixed samples in this method sometimes produce spatially ambiguous and unnatural instances, which can confuse the model when it attempts to locate non-existent spatial information. Therefore, this embodiment introduces an auxiliary contrastive loss, which imposes an additional constraint to maximize the relative similarity (or positive similarity) between the original individual samples and the augmented samples. With the auxiliary contrastive loss, the model is supervised by pre-mixing samples, ensuring that consistent attributes can be extracted between positively mixed individual samples. This supervision provides awareness of the mixed discourse to reduce ambiguity. Furthermore, this embodiment proposes an enhancement function to generate content-rich speech codes.

[0163] The following section provides a further introduction to the learning algorithm for low-resource contrastive learning of speech mixture.

[0164] Figure 9 This is a schematic diagram of a contrastive learning speech fusion model architecture according to an embodiment of this application, such as... Figure 9As shown, the contrastive learning speech mixing model architecture includes a mix-up augmentation module for preprocessing the data and a model training framework that incorporates auxiliary contrastive loss.

[0165] In this embodiment, the hybrid data enhancement module acquires sample 1 and sample 2, and obtains a hybrid sample by performing hybrid enhancement on sample 1 and sample 2.

[0166] Alternatively, common hybrid augmentation typically employs the principle of minimizing neighbor risk to encourage the classifier to perform linearly on training samples. When performing inference on unseen speech data, minimizing neighbor risk reduces undesirable variability. During data preprocessing, two audio instances (sample 1 and sample 2) can be randomly selected from the training set, and a dummy training example can be constructed using the following formula:

[0167]

[0168]

[0169] i, j∈1; N can be used to represent the index of the original training sample set; x and y can represent the original input waveform and one-hot label encoding, respectively; λ~Beta(α, α), for α∈(0,∞) can be the interpolation parameter to determine the amount of content to be linearly mixed, given the virtual input label pair It can be a sample from the training set of the first target sample after hybrid enhancement. The time-spectral characteristics of the short-distance Fourier transform can be used to calculate the cross-entropy (L). mix The losses are as follows:

[0170]

[0171] Where f(.) can be the acoustic coding model, and CE can refer to the standard cross-entropy loss.

[0172] Because highly distorted signals (noise) may be generated from two overlapping speech samples during the mixing and enhancement process, this is unnatural or detrimental to model training. The highly distorted input signal can cause confusion during model training, leading to specific spikes in model error, potentially affecting effective gradients and impairing network convergence. To address this issue, this application proposes a contrastive learning algorithm for speech mixing, introducing an auxiliary component that maximizes the relative similarity between the original mixed samples and the enhanced samples through contrastive learning (auxiliary contrastive loss). This enables the model to generate more effective representations and achieve greater generalization ability under low-resource conditions.

[0173] In this embodiment, the richness of speech-coded content is enhanced by using consistent properties between two enhanced (frontal) views. In this case, the model can be trained using two mixed utterances with the aforementioned training methods, thereby achieving the goal of cultivating a minimal (i.e., low complexity) and sufficient (i.e., high fidelity) embedding.

[0174] Optionally, the above methods may include, such as Figure 9 As shown, randomly selected speech data samples 1 (Xi), 2 (Xj), and a mixture of samples 1 and 2 are obtained. The training process utilizes the three parallel samples mentioned above. Random augmentation can be performed independently on these three parallel samples to obtain augmented samples. The augmented and unaugmented samples can then be passed to a shared encoding network to obtain the latent vector corresponding to each sample. The resulting vectors can be embedded and mapped into a projection dimension of 128. The mixed samples and pre-mixed samples can be compared and learned. Furthermore, predictions can be made on the mixed samples to obtain predicted keywords.

[0175] Alternatively, to calculate the loss for contrastive learning, an L2 norm can be applied to all projected embeddings, and then the mean squared error can be used to measure the similarity between normalized projections. The contrastive loss can be defined by the following formula:

[0176]

[0177] Here, fp(.) is the projection of the model onto r∈{i,j}.

[0178] Optionally, during model training, the training data can be unmixed and mixed samples separately. For example, half of the training data can be unmixed, and half can be mixed, thus achieving a 50% mixing ratio. The training process also includes comparative learning between mixed and premixed samples, and adjusting the model parameters by mixing augmented samples, thereby reducing the variability of perturbation samples and achieving a less complex embedding. The weights relative to the contrastive loss in the target loss function can be calculated using the following formula:

[0179]

[0180] The complete training loss function (target loss function) can be calculated using the following formula:

[0181]

[0182] β can be a penalty parameter that measures the contribution of the contrast loss. It can be determined through testing. For example, it can be set to 0.5. This is just an example and no specific restrictions are placed on the size of the penalty parameter.

[0183] The implementation process of the above method is illustrated below with specific data. It should be noted that the size of the data, the acquisition or processing method are only illustrative examples for the purpose of understanding the scheme, and no specific limitations are imposed here.

[0184] In this embodiment, low-resource training data can be obtained first.

[0185] You can utilize the voice command dataset (V2) from a song, which contains a total of 105,000 utterances, including 35 unique words. Each audio sample in the voice command dataset is stored as a one-second (or shorter) file sampled at 16kHz (e.g., a WAV file).

[0186] For example, a subset of 10 keyword classes can be used for model training, validation, and testing segmentation. These keyword classes can include words such as "up," "down," "left," "right," "yes," "no," "on," "off," "start," and "stop." The speech can be segmented based on the speaker of each word, and the size of the training set can be adjusted accordingly. For instance, experiments can be conducted using training sets of 5%, 10%, 20%, 30%, and 50%, with average training times of 2.5 minutes, 5 minutes, 10 minutes, 15 minutes, and 25 minutes per word, respectively, to obtain the appropriate training data. The training data at these scales is close to the data suitable for training personalized voice wake-up models in practical applications.

[0187] Optionally, by adjusting the speaker partitioning, the diversity of training data can be reduced, the learning difficulty can be increased, and the low-resource model training scenario can be better simulated, so that the trained model can be adapted to a wider range of people in actual deployment.

[0188] In this embodiment, the training data is processed.

[0189] Optionally, the acquired audio samples can be converted into a 64-dimensional log-frequency (log Mel) filter bank (FBank), where the window size of the filter bank can be 25 milliseconds and the offset is 10 milliseconds. The training data can be mixed and augmented; the resolution of the filter bank can be fixed at 98×64 (i.e., equivalent to 1 second of speech), and speech data shorter than 1 second can be padded with zeros to the right.

[0190] Optionally, the training samples can be increased by data fusion augmentation. For example, data in the range of -100ms to 100ms can be randomly time-shifted, and random time stretching with coefficients between 0.9 and 1.1 can be applied. Simultaneously, spectral augmentation methods can be applied, with the masking sizes for time and spectrum set to 13 and 7, respectively.

[0191] In this embodiment, by selecting a suitable neural network model, the model can be trained based on the processed training data.

[0192] Optionally, the voice wake-up model can be a keyword model based on a transformer model (e.g., KWT-1, KWT-3) or a keyword neural visual model based on convolution (e.g., Keyword ConvMixer, ResNet18). These models represent the state-of-the-art in different tuning environments with varying model sizes and complexities. The lightweight ConvMixer model, for example, contains only 0.1M parameters.

[0193] Optionally, a projection can be added to each model, mapping the latent vector embeddings of the training data to a projection dimension of size 128. The projection consists of a linear dense block with a Linear Rectification Function (ReLU) activation. All selected models can be trained in batches of 128. The initial learning is 5e-3, and a step decay rate of 0.85 can be used every four epochs from the 5th to the 70th epoch.

[0194] Alternatively, an optimizer (e.g., Adam) and binary cross-entropy loss can be used during model training.

[0195] To study the quality of acoustic representations using different techniques, in this embodiment of the invention, a visual embedding is achieved on 20% of the training data by using a t-SNE (t-SNE) diagram. Figure 10 This is a schematic diagram of a dimensionality reduction algorithm that compares embeddings of different technologies according to an embodiment of this application, such as... Figure 10 As shown, keyword labels can be set: YES, NO, UP, DOWN, LEFT, RIGHT, ON, OFF, STOP, and GO. The baseline setting attempts to distinguish "RIGHT" and "NO" from other commands; however, the model fails to do so in the rest. Embeddings improve with blending enhancements, and classes become slightly spaced. The class set of this embodiment (Cos-Mix) is the most separable, with only short consonant words grouped together, indicating that this method plays a significant role in learning accurate and content-rich representations.

[0196] Table 1 shows the accuracy of the voice wake-up model across different architectures. As shown in Table 1, the accuracy of keyword classification was determined for each method on training data of different dataset sizes (i.e., 5%, 10%, 20%, 30%, and 50%). When training on small datasets, the performance of all models degrades, with KWT-3 exhibiting the largest performance drop, achieving only 46.5% accuracy on a baseline training dataset of 5%. With less training data (e.g., 5%), the performance gain of the contrastive learning speech mixture method is greater, with KWT-3 showing a relative performance increase of up to 21.7%. Transformer-based models are more prone to performance degradation in low-resource environments due to the lack of training samples for learning highly complex attention mechanisms; however, even in this case, CosMix can help mitigate the low-resource problem through better regularization. Convolutional-based models performed best, with CosMix achieving the highest score of 90% accuracy for keywords on a 5% training set. However, regardless of the model used, the contrastive learning speech mixing method has achieved performance that meets the predetermined conditions on various training set sizes.

[0197] Table 1. Accuracy of the voice wake-up model in different system architectures.

[0198]

[0199]

[0200] Optionally, since the mixing ratio and the parameters of the interpolation weights derived from the β distribution have a significant impact on the performance of the mixing algorithm, a convex curve can be obtained when α, β(α, α) in the β distribution is less than 1, where the amount of audio mixing tends to dominate on one side; but when α is greater than 1, the curve becomes more concave, and the two audio samples are more likely to be mixed proportionally. Table 2 shows the test accuracy based on 20% of the ResNet18 training dataset under different mixing ratios. As shown in Table 2, β(10, 10) generally performs better than β(0.5, 0.5), indicating that for the voice wake-up task, the model can learn more effectively using audio samples mixed in equal proportions. Furthermore, the mixing ratios for the data augmentation algorithm and the contrastive learning speech mixing algorithm are different. The data augmentation algorithm performs best with a mixing ratio of 30%, while the contrastive learning speech mixing algorithm performs best with a mixing ratio of 50%, so a mixing ratio of 50% can be set.

[0201] Table 2 shows the test accuracy on a 20% training dataset based on ResNet18 under different mixing ratios.

[0202]

[0203] This application proposes a contrastive learning speech fusion algorithm, which can enhance the training strategy for speech wake-up models using low-resource training data. This training method utilizes contrastive loss to mitigate the undesirable side effects of noisy training signals generated by traditional fusion training. It can improve model performance under low-resource conditions with various model sizes and can be widely applied to personalized voice wake-up systems for smart devices. It achieves the technical effect of effectively processing training sample sets, solving the technical problem of ineffective training sample set processing.

[0204] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0205] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.

[0206] Example 3

[0207] According to embodiments of this application, a method for implementing the above is also provided. Figure 4 The training sample set generation method shown is a training sample set generation device.

[0208] Figure 11 This is a schematic diagram of a method for generating a training sample set according to an embodiment of this application, as shown below. Figure 11 As shown, the training sample set generation device 1100 may include: a first acquisition unit 1102, a first processing unit 1104, a second processing unit 1106 and a third processing unit 1108.

[0209] The first acquisition unit 1102 is used to acquire the original training sample set to be processed, wherein the original training sample set is used to train the speech processing model.

[0210] The first processing unit 1104 is used to perform hybrid enhancement on the original training sample set to obtain the first target training sample set.

[0211] The second processing unit 1106 is used to perform comparative learning on the first target training sample set and the original training sample set to obtain a comparative loss, wherein the comparative loss is used to characterize the similarity between the original training sample set and the first target training sample set.

[0212] The third processing unit 1108 is used to adjust the first target training sample set based at least on the contrast loss to obtain the second target training sample set, wherein the similarity between the original training sample set and the second target training sample set is greater than the similarity threshold, and the second target training sample set is used to train the speech processing model.

[0213] It should be noted that the first acquisition unit 1102, the first processing unit 1104, the second processing unit 1106, and the third processing unit 1108 mentioned above correspond to steps S402 to S408 in Embodiment 1. The four units and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above units can be hardware or software components stored in memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The above units can also be part of a device and run in the computer terminal A provided in Embodiment 1.

[0214] According to embodiments of this application, a method for implementing the above is also provided. Figure 5 The method for determining the speech processing model shown is a speech processing model determination device.

[0215] Figure 12 This is a schematic diagram of a speech processing model determination device according to an embodiment of this application, such as... Figure 12 As shown, the speech processing model determination device 1200 may include a second acquisition unit 1202 and a training unit 1204.

[0216] The second acquisition unit 1202 is used to acquire a target training sample set, wherein the target training sample set is obtained by adjusting the first target training sample set at least based on the contrast loss, the contrast loss is obtained by comparing and learning the first target training sample set and the original training sample set, and is used to characterize the similarity between the original training sample set and the first target training sample set, and the first target training sample set is obtained by performing hybrid enhancement on the original training sample set.

[0217] Training unit 1204 is used to train a speech processing model based on the target training sample set in response to the similarity between the original training sample set and the target training sample set being greater than a similarity threshold.

[0218] It should be noted that the second acquisition unit 1202 and training unit 1204 mentioned above correspond to steps S502 to S504 in Embodiment 1. The two units and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above units can be hardware or software components stored in memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The above units can also be part of a device and run in the computer terminal A provided in Embodiment 1.

[0219] According to embodiments of this application, a method for implementing the above is also provided. Figure 6The speech processing device of the speech processing method shown.

[0220] Figure 13 This is a schematic diagram of a voice processing device according to an embodiment of this application, such as... Figure 13 As shown, the voice processing device 1300 may include: a collection unit 1302, an extraction unit 1304, and an activation unit 1306.

[0221] The acquisition unit 1302 is used to acquire the voice to be detected sent to the client.

[0222] Extraction unit 1304 is used to extract at least one target keyword from the speech to be detected using a speech processing model. The speech processing model is trained based on a target training sample set. The target training sample set is obtained by adjusting a first target training sample set based on at least a contrast loss. The contrast loss is obtained by comparing and learning the first target training sample set and the original training sample set, and is used to characterize the similarity between the original training sample set and the first target training sample set. The first target training sample set is obtained by performing hybrid enhancement on the original training sample set. The similarity between the original training sample set and the target training sample set is greater than a similarity threshold.

[0223] Activation unit 1306 is used to activate the client based on the target keyword.

[0224] It should be noted that the acquisition unit 1302, extraction unit 1304, and activation unit 1306 mentioned above correspond to steps S602 to S606 in Embodiment 1. The three units and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above units can be hardware or software components stored in memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The above units can also be part of a device and run in the computer terminal A provided in Embodiment 1.

[0225] According to embodiments of this application, a method for implementing the above is also provided. Figure 7 The apparatus shown is for generating training samples.

[0226] Figure 14 This is a schematic diagram of another training sample generation apparatus according to an embodiment of this application, such as... Figure 14 As shown, the training sample generation device 1400 may include: a third acquisition unit 1402, a fourth processing unit 1404, a fifth processing unit 1406, an adjustment unit 1408, and an output unit 1410.

[0227] The third acquisition unit 1402 is used to acquire the original training sample set to be processed by calling the first interface. The original training sample set is used to train the speech processing model. The first interface includes a first parameter, and the parameter value of the first parameter is the original training sample set.

[0228] The fourth processing unit 1404 is used to perform hybrid enhancement on the original training sample set to obtain the first target training sample set.

[0229] The fifth processing unit 1406 is used to perform comparative learning on the first target training sample set and the original training sample set to obtain a comparative loss, wherein the comparative loss is used to characterize the similarity between the original training sample set and the first target training sample set.

[0230] The adjustment unit 1408 is used to adjust the first target training sample set based at least on the contrast loss to obtain the second target training sample set, wherein the similarity between the original training sample set and the second target training sample set is greater than a similarity threshold.

[0231] The output unit 1410 is used to output a second target training sample set by calling a second interface, wherein the output second target training sample set is used to train a speech processing model, and the second interface includes a second parameter, the parameter value of which is the second target training sample set.

[0232] It should be noted that the third acquisition unit 1402, the fourth processing unit 1404, the fifth processing unit 1406, the adjustment unit 1408, and the output unit 1410 mentioned above correspond to steps S702 to S710 in Embodiment 1. The five units and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above units can be hardware or software components stored in memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The above units can also be part of a device and run in the computer terminal A provided in Embodiment 1.

[0233] In the training sample set generation device of this embodiment, the original training sample set is hybridized and enhanced to obtain a first target training sample set. By comparing and learning the first target training sample set and the original training sample set, an auxiliary contrast loss is introduced to maximize the similarity between the original training sample set (original pre-mixed sample set) and the enhanced sample set (first target training sample set). This can increase the speech processing model's cognition of feature content, achieving the goal of improving the performance of the speech processing model under low resource conditions of various sizes. In this way, the technical effect of effectively processing the training sample set is achieved, solving the technical problem of not being able to effectively process the training sample set.

[0234] Example 4

[0235] Embodiments of this application may provide a computer terminal, which may be any computer terminal device in a group of computer terminals. Optionally, in this embodiment, the aforementioned computer terminal may also be replaced by a mobile terminal or other terminal device.

[0236] Optionally, in this embodiment, the computer terminal may be located in at least one of a plurality of network devices in a computer network.

[0237] In this embodiment, the computer terminal described above can execute the program code for the following steps in the method for generating a training sample set of an application: obtaining an original training sample set to be processed, wherein the original training sample set is used to train a speech processing model; performing hybrid enhancement on the original training sample set to obtain a first target training sample set; performing comparative learning on the first target training sample set and the original training sample set to obtain a comparative loss, wherein the comparative loss is used to characterize the similarity between the original training sample set and the first target training sample set; adjusting the first target training sample set based at least on the comparative loss to obtain a second target training sample set, wherein the similarity between the original training sample set and the second target training sample set is greater than a similarity threshold, and the second target training sample set is used to train a speech processing model.

[0238] Optionally, Figure 15 This is a structural block diagram of a computer terminal according to an embodiment of this application. Figure 15 As shown, the computer terminal A may include: one or more (only one is shown in the figure) processors 1502, memory 1504, and transmission devices 1506.

[0239] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the training sample set generation method and apparatus in this application embodiment. The processor executes various functional applications and predictions by running the software programs and modules stored in the memory, thereby realizing the aforementioned training sample set generation method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to computer terminal A via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0240] The processor can invoke information and application programs stored in the memory via a transmission device to perform the following steps: acquiring the original training sample set to be processed, wherein the original training sample set is used to train a speech processing model; performing hybrid enhancement on the original training sample set to obtain a first target training sample set; performing comparative learning on the first target training sample set and the original training sample set to obtain a comparative loss, wherein the comparative loss is used to characterize the similarity between the original training sample set and the first target training sample set; adjusting the first target training sample set based at least on the comparative loss to obtain a second target training sample set, wherein the similarity between the original training sample set and the second target training sample set is greater than a similarity threshold, and the second target training sample set is used to train the speech processing model.

[0241] Optionally, the processor may also execute program code that performs the following steps: mapping a vector composed of speech data from the original training sample set to a first target vector of the target dimension; mapping a vector composed of speech data from the first target training sample to a second target vector of the target dimension, wherein the speech data in the first target training sample is obtained by mixing and enhancing the speech data in the original training sample set; and performing comparative learning on the first target vector and the second target vector to obtain a comparative loss.

[0242] Optionally, the processor may also execute program code that performs the following steps: obtaining the norm value of the first target vector and the norm value of the second target vector, wherein the norm value of the first target vector is used to represent the length of the first target vector and the norm value of the second target vector is used to represent the length of the second target vector; and determining the contrast loss based on the norm value of the first target vector and the norm value of the second target vector.

[0243] Optionally, the processor may also execute program code that performs the following steps: obtaining the mean square error between the second target vector and the first target vector, wherein the mean square error is used to represent the degree of difference between the first target vector and the second target vector; and determining the contrast loss based on the mean square error, the norm of the first target vector, and the norm of the second target vector.

[0244] Optionally, the processor may also execute program code that performs the following steps: in the encoding network model, mapping a vector composed of speech data from the original training sample set to a first target vector of the target dimension, wherein the encoding network model is used to encode the input vector into a vector of the target dimension; mapping a vector composed of speech data from the first target training sample to a second target vector of the target dimension, including: in the encoding network model, mapping a vector composed of speech data from the first target training sample to a second target vector of the target dimension.

[0245] Optionally, the processor may also execute program code that performs the following steps: adjusting the first target training sample set based on cross-entropy loss and contrastive loss to obtain the second target training sample set.

[0246] Optionally, the processor may also execute program code that performs the following steps: adjusting the contrastive loss to the target loss based on the cross-entropy loss; adjusting the first target training sample set based on the target loss to obtain the second target training sample set.

[0247] Optionally, the processor may also execute program code that performs the following steps: comparative learning on the first target training sample set and the enhanced original training sample set to obtain a comparative loss.

[0248] Optionally, the processor may also execute program code that performs random data augmentation on the speech data in the original training sample set.

[0249] Optionally, the processor may also execute program code that performs the following steps: mixing speech data in the original training sample set to obtain mixed speech data; and performing random data augmentation on the mixed speech data to obtain a first target training sample set.

[0250] As an alternative example, the processor can invoke information and applications stored in the memory via a transmission device to perform the following steps: acquiring a target training sample set, wherein the target training sample set is obtained by adjusting a first target training sample set at least based on a contrastive loss, the contrastive loss being obtained by contrastive learning of the first target training sample set and the original training sample set, and used to characterize the similarity between the original training sample set and the first target training sample set, the first target training sample set being obtained by performing hybrid enhancement on the original training sample set; and training a speech processing model based on the target training sample set in response to the similarity between the original training sample set and the target training sample set being greater than a similarity threshold.

[0251] As an alternative example, the processor can invoke information and applications stored in memory via a transmission device to perform the following steps: acquiring the speech to be detected sent to the client; extracting at least one target keyword from the speech to be detected using a speech processing model, wherein the speech processing model is trained based on a target training sample set, the target training sample set is obtained by adjusting a first target training sample set at least based on a contrastive loss, the contrastive loss is obtained by comparing the first target training sample set and the original training sample set, and is used to characterize the similarity between the original training sample set and the first target training sample set, the first target training sample set is obtained by performing hybrid enhancement on the original training sample set, and the similarity between the original training sample set and the target training set is greater than a similarity threshold; activating the client based on the target keyword.

[0252] As an optional example, the processor can invoke information and applications stored in memory via a transmission device to perform the following steps: obtaining a raw training sample set to be processed by invoking a first interface, wherein the raw training sample set is used to train a speech processing model, the first interface includes a first parameter, the parameter value of the first parameter being the raw training sample set; performing hybrid enhancement on the raw training sample set to obtain a first target training sample set; performing comparative learning on the first target training sample set and the raw training sample set to obtain a comparative loss, wherein the comparative loss is used to characterize the similarity between the raw training sample set and the first target training sample set; adjusting the first target training sample set based at least on the comparative loss to obtain a second target training sample set, wherein the similarity between the raw training sample set and the second target training sample set is greater than a similarity threshold; and outputting a second target training sample set by invoking a second interface, wherein the output second target training sample set is used to train a speech processing model, the second interface includes a second parameter, the parameter value of the second parameter being the second target training sample set.

[0253] This application embodiment performs hybrid enhancement on the original training sample set to obtain a first target training sample set. By comparing and learning the first target training sample set and the original training sample set, an auxiliary contrastive loss is introduced to maximize the similarity between the original training sample set (original pre-mixed sample set) and the enhanced sample set (first target training sample set). This can increase the speech processing model's cognition of feature content, achieving the goal of improving the performance of the speech processing model under low resource conditions of various sizes. In this way, it achieves the technical effect of effectively processing the training sample set and solves the technical problem of not being able to effectively process the training sample set.

[0254] Those skilled in the art will understand that Figure 15 The structure shown is for illustrative purposes only. Computer terminal A can also be a smartphone (such as a tablet, a mobile computer, a mobile internet device (MID), a PAD, or other terminal device). Figure 15 This does not limit the structure of the aforementioned computer terminal A. For example, computer terminal A may also include components that are more complex than those described above. Figure 15 Showing more or fewer components (such as network interfaces, display devices, etc.), or having the same Figure 15 The different configurations shown.

[0255] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0256] Example 5

[0257] Embodiments of this application also provide a computer-readable storage medium. Optionally, in this embodiment, the aforementioned computer-readable storage medium can be used to store the program code executed by the training sample set generation method provided in Embodiment 1.

[0258] Optionally, in this embodiment, the computer-readable storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.

[0259] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: obtaining an original training sample set to be processed, wherein the original training sample set is used to train a speech processing model; performing hybrid enhancement on the original training sample set to obtain a first target training sample set; performing comparative learning on the first target training sample set and the original training sample set to obtain a comparative loss, wherein the comparative loss is used to characterize the similarity between the original training sample set and the first target training sample set; adjusting the first target training sample set at least based on the comparative loss to obtain a second target training sample set, wherein the similarity between the original training sample set and the second target training sample set is greater than a similarity threshold, and the second target training sample set is used to train a speech processing model.

[0260] Optionally, the aforementioned computer-readable storage medium may also execute program code that performs the following steps: mapping a vector composed of speech data from the original training sample set to a first target vector of the target dimension; mapping a vector composed of speech data from the first target training sample to a second target vector of the target dimension, wherein the speech data in the first target training sample is obtained by mixing and enhancing the speech data from the original training sample set; and performing comparative learning on the first target vector and the second target vector to obtain a contrastive loss.

[0261] Optionally, the computer-readable storage medium may also execute program code that performs the following steps: obtaining the norm value of a first target vector and the norm value of a second target vector, wherein the norm value of the first target vector is used to represent the length of the first target vector and the norm value of the second target vector is used to represent the length of the second target vector; and determining the contrast loss based on the norm value of the first target vector and the norm value of the second target vector.

[0262] Optionally, the aforementioned computer-readable storage medium may also execute program code that performs the following steps: obtaining the mean square error between the second target vector and the first target vector, wherein the mean square error is used to represent the degree of difference between the first target vector and the second target vector; and determining the contrast loss based on the mean square error, the norm of the first target vector, and the norm of the second target vector.

[0263] Optionally, the aforementioned computer-readable storage medium may also execute program code that performs the following steps: in an encoding network model, mapping a vector composed of speech data from the original training sample set to a first target vector of the target dimension, wherein the encoding network model is used to encode the input vector into a vector of the target dimension; mapping a vector composed of speech data from the first target training sample to a second target vector of the target dimension, including: in an encoding network model, mapping a vector composed of speech data from the first target training sample to a second target vector of the target dimension.

[0264] Optionally, the aforementioned computer-readable storage medium may also execute program code that performs the following steps: adjusting the first target training sample set based on cross-entropy loss and contrastive loss to obtain a second target training sample set.

[0265] Optionally, the aforementioned computer-readable storage medium may also execute program code that performs the following steps: adjusting the contrastive loss to a target loss based on the cross-entropy loss; adjusting the first target training sample set based on the target loss to obtain a second target training sample set.

[0266] Optionally, the aforementioned computer-readable storage medium may also execute program code that performs the following steps: performing comparative learning on the first target training sample set and the enhanced original training sample set to obtain a comparative loss.

[0267] Optionally, the aforementioned computer-readable storage medium may also execute program code that performs random data augmentation on the speech data in the original training sample set.

[0268] Optionally, the aforementioned computer-readable storage medium may also execute program code that performs the following steps: mixing speech data in the original training sample set to obtain mixed speech data; and performing random data augmentation on the mixed speech data to obtain a first target training sample set.

[0269] As an optional example, a computer-readable storage medium is configured to store program code for performing the following steps: obtaining a target training sample set, wherein the target training sample set is obtained by adjusting a first target training sample set at least based on a contrastive loss, the contrastive loss being obtained by contrastive learning of the first target training sample set and the original training sample set, and used to characterize the similarity of the original training sample set relative to the first target training sample set, the first target training sample set being obtained by performing hybrid enhancement on the original training sample set; and training a speech processing model based on the target training sample set in response to the similarity of the original training sample set relative to the target training sample set being greater than a similarity threshold.

[0270] As an optional example, a computer-readable storage medium is configured to store program code for performing the following steps: acquiring speech to be detected sent to a client; extracting at least one target keyword from the speech to be detected using a speech processing model, wherein the speech processing model is trained based on a target training sample set, the target training sample set is obtained by adjusting a first target training sample set at least based on a contrastive loss, the contrastive loss is obtained by contrastive learning of the first target training sample set and the original training sample set, and is used to characterize the similarity of the original training sample set relative to the first target training sample set, the first target training sample set is obtained by hybrid enhancement of the original training sample set, and the similarity of the original training sample set relative to the target training set is greater than a similarity threshold; activating the client based on the target keyword.

[0271] As an optional example, a computer-readable storage medium is configured to store program code for performing the following steps: obtaining a raw training sample set to be processed by calling a first interface, wherein the raw training sample set is used to train a speech processing model, the first interface including a first parameter whose parameter value is the raw training sample set; performing hybrid enhancement on the raw training sample set to obtain a first target training sample set; performing comparative learning on the first target training sample set and the raw training sample set to obtain a comparative loss, wherein the comparative loss is used to characterize the similarity between the raw training sample set and the first target training sample set; adjusting the first target training sample set at least based on the comparative loss to obtain a second target training sample set, wherein the similarity between the raw training sample set and the second target training sample set is greater than a similarity threshold; and outputting a second target training sample set by calling a second interface, wherein the output second target training sample set is used to train a speech processing model, the second interface including a second parameter whose parameter value is the second target training sample set.

[0272] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0273] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0274] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of units or modules may be electrical or other forms.

[0275] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0276] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0277] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0278] The above are merely preferred embodiments of this application. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for generating a training sample set, characterized in that, include: Obtain the original training sample set to be processed, wherein the original training sample set is used to train the speech processing model; The original training sample set is mixed and enhanced to obtain the first target training sample set; A comparative learning process is performed on the first target training sample set and the original training sample set to obtain a comparative loss, wherein the comparative loss is used to characterize the similarity between the original training sample set and the first target training sample set. The first target training sample set is adjusted based at least on the contrast loss to obtain a second target training sample set, wherein the similarity between the original training sample set and the second target training sample set is greater than a similarity threshold, and the second target training sample set is used to train the speech processing model. The process of performing comparative learning on the first target training sample set and the original training sample set to obtain a comparative loss includes: mapping a vector composed of speech data from the original training sample set to a first target vector of the target dimension; mapping a vector composed of speech data from the first target training sample set to a second target vector of the target dimension, wherein the speech data in the first target training sample set is obtained by mixing and enhancing the speech data from the original training sample set; and determining the comparative loss based on the length of the first target vector and the length of the second target vector.

2. The method according to claim 1, characterized in that, The method further includes: Apply the L2 norm to the first target vector and the second target vector to obtain the norm value of the first target vector and the norm value of the second target vector.

3. The method according to claim 2, characterized in that, The method further includes: The norm value of the first target vector is used to represent the length of the first target vector, and the norm value of the second target vector is used to represent the length of the second target vector.

4. The method according to claim 2, characterized in that, The contrast loss is determined based on the norm values ​​of the first target vector and the second target vector, including: The mean square error between the second target vector and the first target vector is obtained, wherein the mean square error is used to represent the degree of difference between the first target vector and the second target vector; The contrast loss is determined based on the mean square error, the norm of the first target vector, and the norm of the second target vector.

5. The method according to claim 1, characterized in that, Mapping the vector composed of speech data from the original training sample set to a first target vector of the target dimension includes: In the encoding network model, a vector composed of speech data from the original training sample set is mapped to the first target vector of the target dimension, wherein the encoding network model is used to encode the input vector into a vector of the target dimension; Mapping a vector composed of speech data from the first target training sample to a second target vector of the target dimension includes: in the encoding network model, mapping a vector composed of speech data from the first target training sample to a second target vector of the target dimension.

6. The method according to claim 1, characterized in that, The method further includes: Determine the cross-entropy loss of the speech data in the original training samples; Adjusting the first target training sample set based at least on the contrastive loss to obtain the second target training sample set includes: adjusting the first target training sample set based on the cross-entropy loss and the contrastive loss to obtain the second target training sample set.

7. The method according to claim 6, characterized in that, Based on the cross-entropy loss and the contrastive loss, the first target training sample set is adjusted to obtain the second target training sample set, including: Based on the cross-entropy loss, the contrastive loss is adjusted to the target loss; The first target training sample set is adjusted based on the target loss to obtain the second target training sample set.

8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: Random data augmentation is performed on the speech data in the original training sample set; The comparative learning process involves comparing the first target training sample set and the original training sample set to obtain a comparative loss, including: comparing the first target training sample set and the enhanced original training sample set to obtain the comparative loss.

9. The method according to any one of claims 1 to 7, characterized in that, The original training samples are mixed and enhanced to obtain a first target training sample set, including: The speech data in the original training sample set are mixed to obtain mixed speech data; Random data augmentation is performed on the mixed speech data to obtain the first target training sample set.

10. The method according to any one of claims 1 to 7, characterized in that, The original training sample set includes labeled data with a data volume less than a data volume threshold.

11. A method for determining a speech processing model, characterized in that, include: A target training sample set is obtained, wherein the target training sample set is obtained by adjusting a first target training sample set based at least on a contrastive loss, the contrastive loss being determined based on the length of the first target vector and the length of the second target vector, and used to characterize the similarity between the original training sample set to be processed and the first target training sample set, the first target vector of the target dimension being obtained by mapping a vector composed of speech data in the original training sample set, the second target vector of the target dimension being obtained by mapping a vector composed of speech data in the first target training sample set, the first target training sample set being obtained by performing hybrid enhancement on the original training sample set, and the speech data in the first target training sample set being obtained by performing hybrid enhancement on the speech data in the original training sample set; In response to the fact that the similarity between the original training sample set and the target training sample set is greater than a similarity threshold, a speech processing model is trained based on the target training sample set.

12. A speech processing method, characterized in that, include: Collect the voice data to be detected sent to the client; At least one target keyword is extracted from the speech to be detected using a speech processing model, wherein the speech processing model is trained based on a target training sample set, the target training sample set is obtained by adjusting a first target training sample set based at least on a contrastive loss, the contrastive loss is determined based on the lengths of the first target vector and the second target vector, and is used to characterize the similarity between the original training sample set to be processed and the first target training sample set, the first target vector of the target dimension is obtained by mapping a vector composed of speech data in the original training sample set, the second target vector of the target dimension is obtained by mapping a vector composed of speech data in the first target training sample set, the first target training sample set is obtained by performing hybrid enhancement on the original training sample set, the speech data in the first target training sample set is obtained by performing hybrid enhancement on the speech data in the original training sample set, and the similarity between the original training sample set and the target training sample set is greater than a similarity threshold; The client is activated based on the target keyword.

13. The method according to claim 12, characterized in that, Activating the client based on the target keyword includes: The client is activated in response to the target keyword having a similarity greater than a predetermined similarity threshold with a predetermined keyword associated with the client.

14. A method for generating a training sample set, characterized in that, include: The original training sample set to be processed is obtained by calling the first interface, wherein the original training sample set is used to train the speech processing model, and the first interface includes a first parameter, the value of which is the original training sample set. The original training sample set is mixed and enhanced to obtain the first target training sample set; A comparative learning process is performed on the first target training sample set and the original training sample set to obtain a comparative loss, wherein the comparative loss is used to characterize the similarity between the original training sample set and the first target training sample set. The first target training sample set is adjusted based at least on the contrast loss to obtain a second target training sample set, wherein the similarity between the original training sample set and the second target training sample set is greater than a similarity threshold. The second target training sample set is output by calling the second interface, wherein the output second target training sample set is used to train the speech processing model, and the second interface includes a second parameter, the parameter value of which is the second target training sample set; The process of performing comparative learning on the first target training sample set and the original training sample set to obtain a comparative loss includes: mapping a vector composed of speech data from the original training sample set to a first target vector of the target dimension; mapping a vector composed of speech data from the first target training sample set to a second target vector of the target dimension, wherein the speech data in the first target training sample set is obtained by mixing and enhancing the speech data from the original training sample set; and determining the comparative loss based on the length of the first target vector and the length of the second target vector.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein when the program is run by a processor, it controls the device in which the computer-readable storage medium resides to perform the method according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Training sample generation method and device, computer equipment and storage medium

    CN113361629A

  • Audio recognition method and device, computer equipment and computer readable storage medium

    CN115359785A