Audio processing method and apparatus, device and storage medium
By acquiring the feature representation of audio clips and using machine learning models for end-to-end entity recognition, the problems of information loss and recognition error in the speech recognition process in the prior art are solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- PCT/CN2024/140073
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-18
- Filing Date
- 2024-12-17
- Publication Date
- 2025-06-26
AI Technical Summary
The existing audio processing technology has information loss and recognition errors in the speech recognition process, and the phased model training is not direct enough, resulting in insufficient recognition accuracy and robustness.
By obtaining feature representations of multiple audio clips in the target speech, end-to-end entity recognition is used to generate a sequence of target text associated with the target speech, including entity words and their types.
It improves the accuracy and robustness of speech recognition, reduces information loss and recognition errors, and realizes direct recognition of entity information in the target speech.
Smart Images

Figure CN2024140073_26062025_PF_FP_ABST
Abstract
Description
Method, apparatus, device and storage medium for audio processing
[0001] This application claims priority to Chinese patent application number 2023117459108, filed on December 18, 2023, and entitled “Methods, devices, apparatus and storage media for audio processing”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] Example embodiments of the present disclosure relate generally to the field of computers, and more particularly, to methods, apparatuses, devices, and computer-readable storage media for audio processing. Background Art
[0003] With the development of computer technology, many model-based auxiliary functions and tools have emerged. These auxiliary functions or tools (such as smart home systems, voice assistants, and intelligent robots) can solve problems in people's daily lives and work through technologies such as machine learning, natural language processing, and image recognition. As a result, they can improve people's efficiency and convenience and reduce their work pressure. Summary of the Invention
[0004] In a first aspect of the present disclosure, a method for audio processing is provided. The method comprises: obtaining multiple audio feature representations corresponding to multiple audio segments in a target speech, the multiple audio segments including at least a first audio segment and a second audio segment, the first audio segment being a preceding audio segment of the second audio segment; and utilizing a trained machine learning model to perform the following operations: determining a first text based on a first audio feature representation corresponding to the first audio segment; extracting a first text feature representation of the first text; determining a second text based on a second audio feature representation and the first text feature representation corresponding to the second audio segment; and determining a target text sequence associated with the target speech based at least on the first text and the second text, the target text sequence comprising at least one entity word and a corresponding entity type appearing in the target speech.
[0005] In a second aspect of the present disclosure, a device for audio processing is provided. The device includes: an audio feature representation acquisition module configured to acquire multiple audio feature representations corresponding to multiple audio segments in a target speech, the multiple audio segments including at least a first audio segment and a second audio segment, the first audio segment being a preceding audio segment of the second audio segment; and a text sequence determination module configured to use a trained machine learning model to perform the following operations: determine a first text based on a first audio feature representation corresponding to the first audio segment; extract a first text feature representation of the first text; determine a second text based on a second audio feature representation and the first text feature representation corresponding to the second audio segment; and determine a target text sequence associated with the target speech based at least on the first text and the second text, the target text sequence including at least one entity word and a corresponding entity type appearing in the target speech.
[0006] In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method of the first aspect of the present disclosure.
[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium and can be executed by a processor to perform the method according to the first aspect of the present disclosure.
[0008] It should be understood that the contents described in the content of this disclosure are not intended to limit the key features or important features of the embodiments of this disclosure, nor are they intended to limit the scope of this disclosure. Other features of this disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent hereinafter with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0010] FIG1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;
[0011] FIG2 shows a flowchart of a process for audio processing according to some embodiments of the present disclosure;
[0012] FIG3 shows a schematic diagram of an example architecture of a machine learning model according to some embodiments of the present disclosure;
[0013] FIG4 is a schematic diagram showing an example of a training sample set according to some embodiments of the present disclosure;
[0014] FIG5 shows a schematic diagram of an example of applying a machine learning model according to some embodiments of the present disclosure;
[0015] FIG6 shows a block diagram of an apparatus for audio processing according to some embodiments of the present disclosure; and
[0016] FIG7 illustrates a block diagram of an electronic device in which one or more embodiments of the present disclosure may be implemented. DETAILED DESCRIPTION
[0017] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Instead, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0018] It should be noted that the acquisition, storage and application of user personal information involved in the technical solution of this disclosure are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0019] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to." The term "based on" should be understood as "based at least in part on." The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment." The term "some embodiments" should be understood as "at least some embodiments." Other explicit and implicit definitions may be included below.
[0020] As briefly mentioned above, intelligent assistance can improve people's efficiency and convenience and reduce their work pressure. For example, in call center systems, there are many auxiliary functions, such as intent recognition, entity recognition, and group complaint event judgment. These auxiliary functions can help customer service staff handle business better. In the voice scenario, some solutions are based on the architecture of user voice to Automatic Speech Recognition (ASR) system and then to Natural Language Processing (NLP) system. That is, the user voice is transcribed into corresponding text by the ASR system, and then the NLP system recognizes intent and entities based on the text.
[0021] However, this approach presents some challenges. For example, some speech may be lost as it passes through the ASR system, preventing downstream systems from using richer speech information. Another example is that the ASR system's recognition errors can be transmitted to downstream systems. This error transmission can lead to erroneous recognition results. Furthermore, in certain scenarios, some redundant information does not need to be recognized, and further transcription may lead to errors in the final recognition results. In some specific application scenarios, such as call center systems providing flight booking services, the key information is the user's departure, destination, and date; all other user information can be temporarily ignored.
[0022] Furthermore, the above approach divides the recognition task into two phases, with each phase performing model training based on its own data. Firstly, this approach is not straightforward in terms of results. Secondly, making the final model more robust also requires a significant amount of data.
[0023] In order to at least partially solve the above problems, the present disclosure provides an improved solution for audio processing. According to the solution, a device obtains multiple audio feature representations corresponding to multiple audio segments in the target speech. The multiple audio segments include at least a first audio segment and a second audio segment, and the first audio segment is a preceding audio segment of the second audio segment. The device determines a first text based on the first audio feature representation corresponding to the first audio segment using a trained machine learning model. Further, the device extracts a first text feature representation of the first text. The device determines a second text based on the second audio feature representation and the first text feature representation corresponding to the second audio segment. The device determines a target text sequence associated with the target speech based on at least the first text and the second text. Such a target text sequence includes at least one entity word and a corresponding entity type that appear in the target speech. Thus, the corresponding entity result can be generated directly through speech to achieve end-to-end entity recognition. In this way, recognition accuracy can be improved.
[0024] Hereinafter, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings.
[0025] FIG1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In environment 100, a terminal device 110 or its attached device can capture or store user speech (also referred to as target speech). In some application scenarios, such as a call center system providing ticket reservation services, the terminal device 110 or its attached device can capture user speech in real time.
[0026] In some embodiments of the present disclosure, the machine learning model 130 used can be deployed on the remote device 120. The terminal device 110 can communicate with the remote device 120 (for example, via network communication) to utilize the machine learning model 130 stored thereon to perform model reasoning tasks, that is, audio processing tasks for the target voice. In such a case, the target voice can be sent to the remote device 120 by the terminal device 110. In some embodiments, although not shown in the figure, the machine learning model 130 can also be partially or completely deployed locally on the terminal device 110 and run locally by the terminal device 110 to perform audio processing tasks for the target voice. The embodiments of the present disclosure do not impose specific restrictions on this.
[0027] The terminal device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. The remote device 120 may, for example, include a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, a virtual machine, and the like. Although a single device is shown, the remote device 120 may include multiple physical devices. In addition, although only a single terminal device 110 is shown, the remote device 120 or the machine learning model 130 deployed therein can be accessed by multiple terminal devices 110 to provide reasoning capabilities of the machine learning model 130.
[0028] It should be understood that the structure and functionality of environment 100 are described for exemplary purposes only and do not imply any limitation on the scope of the present disclosure.
[0029] FIG2 shows a flowchart of a process 200 for audio processing according to some embodiments of the present disclosure. Process 200 may be implemented at the terminal device 110 and / or the remote device 120. For ease of description, the following description will be based on an example of implementation at the terminal device 110 and with reference to FIG1 .
[0030] At block 210, the terminal device 110 obtains multiple audio feature representations corresponding to multiple audio segments in the target speech. For example, the terminal device 110 includes an audio acquisition device. The audio acquisition device can capture the target speech (e.g., the speech of a user after accessing a call center system). In some embodiments, the terminal device 110 can utilize the trained machine learning model 130 to divide the target speech into multiple audio segments and extract an audio feature representation from each audio segment, thereby obtaining multiple audio feature representations corresponding to the multiple audio segments.
[0031] The multiple audio segments include at least a first audio segment and a second audio segment. The first audio segment is a preceding audio segment of the second audio segment. For example, the terminal device 110 acquires the target speech and divides the target speech into audio segments 1, audio segments 2, audio segments 3, and so on in chronological order. Audio segment 1 can be called the first audio segment, and audio segment 2 can be called the second audio segment. Similarly, if audio segment 2 is called the first audio segment, then audio segment 3 can be called the second audio segment.
[0032] At block 220, the terminal device 110 performs the audio processing task using the trained machine learning model 130. Specifically, the machine learning model 130 receives as input multiple audio feature representations corresponding to multiple audio segments in the target speech and outputs a target text sequence associated with the target speech. This target text sequence includes entity words and corresponding entity types that appear in the target speech. This allows redundant information in the target speech to be removed, allowing only the entity information in the target speech to be recognized.
[0033] Specifically, the trained machine learning model 130 receives a first audio feature representation corresponding to a first audio segment and outputs a first text. Further, the trained machine learning model 130 receives as input a first text and a second audio feature representation identified from the first audio segment. The trained machine learning model 130 extracts a first text feature representation from the first text, and determines a second text based on the second audio feature representation and the first text feature representation corresponding to the second audio segment. That is, for text recognition of a subsequent audio segment (i.e., the second audio segment), the corresponding text recognition result can be determined based on the audio feature information of the audio segment itself and in combination with the text feature information of its preceding audio segment. Furthermore, the trained machine learning model 130 determines a target text sequence based at least on the first text and the second text.
[0034] In some embodiments, the target speech further includes more audio segments, for example, a third audio segment. The trained machine learning model 130 then receives the second text and a third audio feature representation corresponding to the third audio segment as input, and extracts a second text feature representation from the second text. Furthermore, the trained machine learning model 130 determines a third text based on the third audio feature representation and the second text feature representation, and determines a target text sequence based at least on the first text, the second text, and the third text.
[0035] Thus, the entity information contained in the previous audio segment can be used to assist in confirming the entity information contained in the current audio segment. In this way, the accuracy of the recognition result can be improved.
[0036] The following will refer to Figures 3 and 4 to describe the training process of the machine learning model 130, and will refer to Figure 5 to describe the application process of the machine learning model 130.
[0037] FIG3 illustrates a schematic diagram of an example architecture 300 of a machine learning model 130 according to some embodiments of the present disclosure. The architecture 300 generally includes an audio information encoding network 310, a text information encoding network 320, and a joint network 330. The architecture 300 can be implemented at the terminal device 110 and / or the remote device 120. For ease of description, the following description assumes implementation at the terminal device 110 and reference to FIG1 .
[0038] In some embodiments, the training data for the machine learning model 130 may include a speech sample set and a text sequence sample set. The speech sample set includes at least one sample speech 301. The text sequence sample set includes at least one sample text sequence 303. The sample text sequence 303 may include multiple entity words and corresponding entity types associated with the sample speech 301. Entity words may include words, phrases, entity names, etc. with specific meanings or categories. Entity types include, for example, names of people, places, names of organizations, dates, times, etc. Categorizing and labeling different types of entity words helps improve the accuracy and efficiency of reasoning tasks.
[0039] In some embodiments, the architecture 300 can be a sequence-to-sequence (seq2seq) structure, for example, processing sequence-to-sequence tasks through an encoder and a decoder. For example, the audio information encoding network 310 can be used as an encoder to extract audio feature representation 302 from the sample speech 301. The text information encoding network 320 can be used as an encoder to extract text feature representation 304 from the sample text sequence 303. The joint network 330 acts as a decoder, receiving the audio feature representation 302 and the text feature representation 304 to generate a predicted text sequence 305. Such a predicted text sequence 305 can be a fusion of the audio feature representation corresponding to the current speech segment and the text feature representation corresponding to the previous speech segment, and outputting a probability distribution of the text sequence.
[0040] The network structures of the audio information encoding network 310, the text information encoding network 320, and the joint network 330 include, but are not limited to, a recurrent neural network (RNN), a long short-term memory (LSTM) network, a gated recurrent unit (GRU), a transformer, etc. This disclosure does not impose any restrictions on this.
[0041] The above describes the architecture of the machine learning model 130. Before model training, training data needs to be constructed first. In some specific scenarios, such as a call center system that provides ticket booking services, it is only necessary to train the machine learning model 130 to identify key information in the user's voice. In some embodiments, such key information includes entity information, such as entity words and corresponding entity types. Such entity types may include at least one of departure, destination, date, and time. Therefore, the training goal of the machine learning model 130 is to predict the entity information in the target voice. The required samples mainly involve voice, entity words, and entity types, while other information can be directly ignored. The following will take the scenario of a call center system that provides ticket booking services as an example and refer to Figure 4 to describe an example of a training sample set.
[0042] Figure 4 shows a schematic diagram of multiple examples of training sample sets according to some embodiments of the present disclosure. The sample speech may include historical speech data collected between customer service and users. The sample text sequence may include the entity content in the work order filled out by the customer service.
[0043] In some embodiments, the sample text sequence includes multiple entity words and corresponding entity types associated with the sample speech, and the multiple entity words and corresponding entity types are arranged in the order of appearance in the associated speech sample. For example, in the example of Figure 4, the text 404 corresponding to the sample speech 402 includes "I will fly from City A to City B tomorrow", and the constructed sample text sequence 406 may include the entity word "tomorrow" and its entity type "date", the entity word "City A" and its entity type "departure place", and the entity word "City B" and its entity type "destination" arranged in sequence.
[0044] Since manual customer service can directly determine the entity information, the text sequence output by the machine learning model 130 only needs to include the correct entity set. Therefore, the order of entity words and corresponding entity types can be disrupted when constructing training data. In some embodiments, the multiple entity words and corresponding entity types associated with each voice sample can be randomly arranged. For example, the text 404 corresponding to the sample voice 402 includes "I will fly from City A to City B tomorrow", then the constructed sample text sequence 408 can include the randomly arranged entity word "City B" and its entity type "destination", the entity word "tomorrow" and its entity type "date", the entity word "City A" and its entity type "departure place". In this way, more training data can be constructed to improve the robustness of the machine learning model 130. In some complex cases, even with the influence of factors such as background music, the trained machine learning model 130 can accurately identify the entity set.
[0045] The method of disrupting the order of entities contained in the text sequence sample can also be referred to as performing entity data enhancement on the text sequence sample set. In some embodiments, training the machine learning model 130 can also include performing voice data enhancement on the speech sample set. Such voice data enhancement can include one or more of time stretching, pitch shifting, and spectrum expansion.
[0046] Time stretching involves shortening or lengthening the duration of a speech sample while maintaining the audio's frequency spectrum. This allows the duration of a speech sample to be altered without affecting its tone or content. This improves the machine learning model's 130 adaptability to different speakers, ambient noise, and other conditions.
[0047] Pitch shifting involves increasing or decreasing the pitch of a sample speech while maintaining the same speaking rate. In this way, the adaptability of the machine learning model 130 to various speakers and differences in emotional expression can be improved.
[0048] Spectrum expansion includes using the spectrogram to perform random blocking, masking, or distortion on the time axis and the spectrum axis. In this way, the adaptability of the machine learning model 130 to speech representation under different noise conditions can be improved.
[0049] The above describes how to construct training data. The following will describe how to apply the trained machine learning model 130 with reference to Figure 5.
[0050] 5 shows a schematic diagram of an example process 500 of applying a machine learning model 130 according to some embodiments of the present disclosure. The target speech, for example, includes at least a first audio segment 501 and a second audio segment 504. The process 500, for example, includes a first stage and a second stage.
[0051] In the first stage, the speech information encoding network 310 extracts a first audio feature representation 502 from a first audio segment 501. The joint network 330 outputs a first text 503 based on the first audio feature representation 502. Such first text 503 will serve as the input of the second stage.
[0052] In the second stage, the speech information encoding network 310 extracts a second audio feature representation 505 from the second audio segment 504. The text information encoding network 320 receives the first text 503 output by the joint network 330 in the first stage and extracts a first text feature representation 506 from the first text 503. The joint network 330 fuses the second audio feature representation 505 and the first text feature representation 506 and outputs a second text 507.
[0053] Further, the machine learning model 130 determines a target text sequence 508 based on the first text 503 and the second text 507 .
[0054] If the target speech includes more audio segments, process 500 may include more stages. For example, if the target speech includes a third audio segment, the third stage of process 500 includes extracting a second text feature representation from the second text output by joint network 330 in the second stage, and determining a third text based on the third audio feature representation and the second text feature representation extracted from the third audio segment. Finally, machine learning model 130 determines a target text sequence 508 based on first text 503, second text 507, and third text.
[0055] In some embodiments, if the first audio segment is at the beginning of the target speech, the first text may include a start symbol. The start symbol may include, for example, one or more characters such as letters, punctuation marks, special symbols, etc. Such a start symbol may indicate the beginning of the target speech. For example, adding the character [start] at the beginning of the first text may provide a clear start marker for the machine learning model 130, so that the machine learning model 130 may generate a text sequence starting from this character.
[0056] In some embodiments, if the second audio segment is at the end of the target speech, the second text may include an end symbol. The end symbol may include, for example, one or more characters such as letters, punctuation marks, special symbols, etc. Such an end symbol may indicate the end of the target speech. For example, adding the [end] character at the end of the second text may provide a clear end marker for the machine learning model 130, allowing the machine learning model 130 to determine the length of the text sequence, thereby avoiding infinite loops or incomplete output.
[0057] As an example, in a call center system, a customer service representative listens to a user's speech. Machine learning model 130 receives the user's speech and generates each entity and its corresponding entity type, starting with the character [start] and ending with the character [end], outputting the corresponding entity set. The customer service representative can then select and fill in the slots in the output entity set, thereby completing the flight reservation for the user.
[0058] In summary, this paper proposes an end-to-end entity set recognition solution that uses a trained machine learning model to generate corresponding entity results from speech, ignoring intermediate losses. During the model training process, the speech and corresponding entity set results are directly used for training. In addition, the use of speech data augmentation and entity data augmentation increases the diversity of the training data, thereby improving the robustness of the model.
[0059] FIG6 shows a block diagram of an apparatus 600 for audio processing according to some embodiments of the present disclosure. Apparatus 600 may be implemented in terminal device 110 and / or remote device 120 of FIG1 . Each module / component in apparatus 600 may be implemented by hardware, software, firmware, or any combination thereof.
[0060] The device 600 includes an audio feature representation acquisition module 610, which is configured to obtain multiple audio feature representations corresponding to multiple audio segments in the target speech, the multiple audio segments including at least a first audio segment and a second audio segment, the first audio segment being a preceding audio segment of the second audio segment. The device 600 also includes a text sequence determination module 620, which is configured to use a trained machine learning model to perform the following operations: determine a first text based on a first audio feature representation corresponding to the first audio segment; extract a first text feature representation of the first text; determine a second text based on a second audio feature representation and the first text feature representation corresponding to the second audio segment; and determine a target text sequence associated with the target speech based on at least the first text and the second text, the target text sequence including at least one entity word and a corresponding entity type appearing in the target speech.
[0061] In some embodiments, the training data of the machine learning model includes a speech sample set and a text sequence sample set, and the text sequence sample set includes multiple entity words and corresponding entity types associated with each speech sample in the speech sample set.
[0062] In some embodiments, the plurality of entity words and corresponding entity types associated with each voice sample are arranged in an order of appearance in the associated voice sample.
[0063] In some embodiments, multiple entity words and corresponding entity types associated with each speech sample are randomly arranged.
[0064] In some embodiments, training of the machine learning model includes performing speech data enhancement on a speech sample set, where the speech data enhancement includes one or more of time stretching, pitch shifting, and spectrum expansion.
[0065] In some embodiments, the text sequence determination module 620 is configured to determine the first text to include at least a start symbol indicating the start of the target speech if the first audio segment is located at the start of the target speech.
[0066] In some embodiments, the text sequence determination module 620 is configured to determine the second text to include at least an end symbol indicating the end of the target speech if the second audio segment is located at the end of the target speech.
[0067] The units included in the device 600 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the units in the device 600 can be implemented at least in part by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0068] FIG7 shows a block diagram of an electronic device 700 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 700 shown in FIG7 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 700 shown in FIG7 can be used to implement the terminal device 110 and / or the remote device 120 of FIG1 .
[0069] As shown in FIG7 , electronic device 700 is a general-purpose electronic device. Components of electronic device 700 may include, but are not limited to, one or more processors or processing units 710, memory 720, storage device 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. Processing unit 710 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 720. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing capabilities of electronic device 700.
[0070] The electronic device 700 typically includes a plurality of computer storage media. Such media can be any available media accessible to the electronic device 700, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 720 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory) or some combination thereof. The storage device 730 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk or any other medium, which can be used to store information and / or data (e.g., training data for training) and can be accessed within the electronic device 700.
[0071] The electronic device 700 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 7 , a disk drive for reading from or writing to a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 720 may include a computer program product 725 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
[0072] The communication unit 740 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 700 can be implemented as a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 700 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.
[0073] Input device 750 may be one or more input devices, such as a mouse, keyboard, or trackball. Output device 760 may be one or more output devices, such as a display, a speaker, or a printer. Electronic device 700 may also communicate with one or more external devices (not shown) via communication unit 740 as needed, such as storage devices, display devices, or the like, with one or more devices that allow a user to interact with electronic device 700, or with any device that allows electronic device 700 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0074] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which one or more computer instructions are stored, wherein the one or more computer instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.
[0075] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0076] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable audio processing device, thereby producing a machine such that, when these instructions are executed by the processing unit of the computer or other programmable audio processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowcharts and / or block diagrams. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable audio processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowcharts and / or block diagrams.
[0077] Computer-readable program instructions may also be loaded onto a computer, other programmable audio processing device, or other device, so that a series of operational steps are performed on the computer, other programmable audio processing device, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable audio processing device, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0078] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.
[0079] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, non-exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the implementations disclosed herein.
Claims
1. An audio processing method, comprising: Acquire multiple audio feature representations corresponding to multiple audio segments in the target speech, where the multiple audio segments include at least a first audio segment and a second audio segment, where the first audio segment is a preceding audio segment of the second audio segment; and Using the trained machine learning model, do the following: Determining a first text based on a first audio feature representation corresponding to the first audio segment; extracting a first text feature representation of the first text; Determining a second text based on a second audio feature representation corresponding to the second audio segment and the first text feature representation; as well as Based at least on the first text and the second text, a target text sequence associated with the target speech is determined, the target text sequence including at least one entity word and a corresponding entity type appearing in the target speech.
2. The method according to claim 1, wherein the training data of the machine learning model includes a speech sample set and a text sequence sample set, and the text sequence sample set includes multiple entity words and corresponding entity types associated with each speech sample in the speech sample set.
3. The method according to claim 2, wherein the multiple entity words and corresponding entity types associated with each voice sample are arranged in the order of appearance in the associated voice sample.
4. The method according to claim 2, wherein multiple entity words and corresponding entity types associated with each speech sample are randomly arranged.
5. The method according to claim 2, wherein the training of the machine learning model includes performing speech data enhancement on the speech sample set, and the speech data enhancement includes one or more of time stretching, pitch shifting, and spectrum expansion.
6. The method of claim 1, wherein based on the first audio feature representation, determining the first text comprises: If the first audio segment is located at the beginning of the target speech, the first text is determined to include at least a start symbol indicating the beginning of the target speech.
7. The method of claim 1 , wherein determining the second text based on the second audio feature representation and the first text feature representation comprises: If the second audio segment is located at the end of the target speech, the second text is determined to include at least an end symbol indicating the end of the target speech. The method according to claim 1 , wherein the entity type comprises at least one of a departure place, a destination, a date, and a time.
9. An apparatus for audio processing, comprising: an audio feature representation acquisition module, configured to acquire a plurality of audio feature representations corresponding to a plurality of audio segments in a target speech, wherein the plurality of audio segments at least include a first audio segment and a second audio segment, wherein the first audio segment is a preceding audio segment of the second audio segment; and The text sequence determination module is configured to use the trained machine learning model to perform the following operations: Determining a first text based on a first audio feature representation corresponding to the first audio segment; extracting a first text feature representation of the first text; Determining a second text based on a second audio feature representation corresponding to the second audio segment and the first text feature representation; as well as Based at least on the first text and the second text, a target text sequence associated with the target speech is determined, the target text sequence including at least one entity word and a corresponding entity type appearing in the target speech.
10. An electronic device, comprising: at least one processing unit; as well as At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the device to perform the method according to any one of claims 1 to 8 when executed by the at least one processing unit.
11. A computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the method according to any one of claims 1 to 8 when executed by a processor.
Citation Information
Patent Citations
Event detection method and system based on audio
CN111863029A
Voice recognition model training method and device, voice recognition method and device, equipment and medium
CN111951789A
Voice recognition method and device, electronic equipment and medium
CN112530408A
Interaction method, device and system and intelligent equipment
CN113836932A
Method and device for training speech recognition model, electronic equipment and storage medium
CN113889088A
Cited By
Method and device for generating audio content, equipment, storage medium and program product
CN120913537A