Audio processing method, apparatus, system, storage medium, and electronic device
By acquiring and aligning audio data, identifying and mixing audio data, the system can recognize and distinguish the speaking texts of participants, solving the problem of difficulty in understanding meeting content caused by poor recording quality, and achieving more accurate understanding of meeting content.
Patent Information
- Application Number
- CN202311031513.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-15
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2043-08-15
AI Technical Summary
In audio conferences, poor recording quality or unclear speech can make it difficult for users to accurately understand the meeting content.
By acquiring the audio data, participant identifiers, time stamps, and mixed audio data of each participant, the speech text is identified and aligned to distinguish the speech content of each participant, and mixed speech text data associated with the participant identifier is generated.
It effectively distinguishes the content of each participant's speech, helping users to understand the meeting content more accurately.
Smart Images

Figure CN119495301B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio conferencing, and in particular to an audio processing method, apparatus, system, storage medium, and electronic device. Background Technology
[0002] With the development of the internet, more and more users are holding audio and video conferences online. In these conferences, it's often necessary to record the participants' statements. The common technology for this is to record the audio of the meeting, allowing users to access the content from the recorded audio file.
[0003] However, due to poor recording quality or unclear audio from the speaker, users cannot accurately understand the meeting content through the audio files, thus affecting their comprehension of the meeting content. Summary of the Invention
[0004] To overcome the problems existing in related technologies, this application provides an audio processing method, apparatus, system, storage medium, and electronic device that effectively distinguishes the speech content of each participant, making it easier for users to understand the meeting content.
[0005] According to a first aspect of the embodiments of this application, an audio processing method is provided, comprising the following steps:
[0006] The system acquires audio data from each participant, participant identifiers corresponding to each audio data, first time identifiers for each audio data, mixed audio data, and second time identifiers for the mixed audio data; the mixed audio data is obtained by mixing the audio data from each participant.
[0007] Identify each of the audio data points to obtain the spoken text data;
[0008] Based on each of the first time identifiers and the second time identifiers, each of the participant identifiers, the speech text data, and the mixed audio data are aligned to obtain the mixed speech text data associated with the participant identifier.
[0009] According to a second aspect of the embodiments of this application, an audio processing apparatus is provided, comprising:
[0010] The data acquisition module is used to acquire the audio data of each participant, the participant identifier corresponding to each audio data, the first time identifier of each audio data, the mixed audio data, and the second time identifier of the mixed audio data; the mixed audio data is obtained by mixing the audio data of each participant.
[0011] The text recognition module is used to recognize each audio data item and obtain the spoken text data;
[0012] The speech text acquisition module is used to align the participant identifiers, the speech text data and the mixed audio data according to the first time identifier and the second time identifier to obtain the mixed speech text data associated with the participant identifiers.
[0013] According to a third aspect of the embodiments of this application, an audio processing system is provided, comprising: a server and a plurality of participating terminals; each of the participating terminals is connected to the server;
[0014] Each participating terminal sends its own audio data and the participant identifier corresponding to each audio data to the server;
[0015] The server is used to acquire audio data from each participant, participant identifiers corresponding to each audio data, first time identifiers of each audio data, mixed audio data, and second time identifiers of the mixed audio data; the mixed audio data is obtained by mixing the audio data from each participant.
[0016] The server is used to identify each audio data item and obtain spoken text data;
[0017] The server is used to align each participant identifier, the speech text data, and the mixed audio data according to each of the first time identifier and the second time identifier, so as to obtain the mixed speech text data associated with the participant identifier.
[0018] According to a fourth aspect of the present application, an electronic device is provided, including a processor and a memory; the memory stores a computer program adapted to be loaded by the processor and executed as described above in the audio processing method.
[0019] According to a fifth aspect of the embodiments of this application, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the audio processing method as described above.
[0020] This application embodiment acquires audio data of each participant, participant identifiers corresponding to each audio data, first time identifiers of each audio data, mixed audio data, and second time identifiers of the mixed audio data; the mixed audio data is obtained by mixing the audio data of each participant; each audio data is identified to obtain speech text data; based on each first time identifier and second time identifier, each participant identifier, the speech text data, and the mixed audio data are aligned to obtain mixed speech text data associated with the participant identifier, thereby effectively distinguishing the speech content of each participant based on the mixed speech text data, making it easier for users to understand the meeting content.
[0021] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application.
[0022] To better understand and implement this invention, the following detailed description is provided in conjunction with the accompanying drawings. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a schematic diagram illustrating an application scenario of an audio processing method according to one embodiment of this application;
[0025] Figure 2 This is a flowchart illustrating an audio processing method according to one embodiment of this application;
[0026] Figure 3 This is a flowchart illustrating a method for determining the audio data of each participant and the participant identifier corresponding to each audio data, as shown in one embodiment of this application.
[0027] Figure 4 A flowchart illustrating a method for determining a first time identifier is shown in one embodiment of this application;
[0028] Figure 5 A flowchart illustrating a method for determining a second time identifier is shown in one embodiment of this application;
[0029] Figure 6 A flowchart illustrating a method for determining mixed-stream speech text data according to an embodiment of this application;
[0030] Figure 7 This is a schematic diagram illustrating the principle of determining mixed-stream speech text data in one embodiment of this application;
[0031] Figure 8 A flowchart illustrating an audio processing method according to another embodiment of this application;
[0032] Figure 9 This is a schematic diagram illustrating the effect of an audio processing method according to one embodiment of this application;
[0033] Figure 10 This is a schematic block diagram of an audio processing apparatus according to one embodiment of this application;
[0034] Figure 11This is a schematic diagram of the structure of an electronic device according to one embodiment of this application. Detailed Implementation
[0035] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings. Wherein, when the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements.
[0036] It should be understood that the embodiments described below do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this application.
[0037] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms "a" and "the" as used herein are also intended to include the plural forms unless the context clearly indicates otherwise. Furthermore, in the description of this application, unless otherwise stated, "a plurality of" means two or more. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more associated listed items, for example, A and / or B, which can represent: A alone, A and B together, and B alone; the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0038] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, this information should not be limited to these terms, and these terms are only used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence, nor should they be construed as indicating or implying relative importance. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances. Depending on the context, the word "if" as used in this application can be interpreted as "when," "when," or "in response to determination."
[0039] It should be understood that although the steps in this application are identified by their sequence number and displayed sequentially in the method flowchart according to the arrows, these steps are not necessarily executed in the order indicated by the step sequence number and the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Each step may include multiple sub-steps or multiple stages, which are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.
[0040] The audio processing method of this application is mainly applicable to scenarios involving multi-terminal audio interaction, such as audio conferencing scenarios, audio-video conferencing scenarios, and remote classroom teaching scenarios. This application provides an example illustration using an audio conferencing scenario.
[0041] Please see Figure 1 This is a schematic diagram illustrating an application scenario of an audio processing method according to an embodiment of this application. The application scenario includes a server 110 and several participating terminals 120; the several participating terminals 120 interact with each other through the server 110.
[0042] In this application, several participating terminals 120 can access the Internet via a network to establish a data transmission channel with the server 110, thereby enabling audio data interaction. The network can be any type of communication medium capable of enabling communication between the participating terminals 120 and the server 110, such as a wired communication link, a wireless communication link, or a fiber optic cable, etc., and this application does not impose any limitations on this.
[0043] Server 110 can serve as a business server, responsible for further connecting to relevant audio data servers, video streaming servers, and other servers providing related support, thereby forming a logically interconnected service cluster to provide services to relevant terminal devices, such as... Figure 1 The participating terminal 120 shown provides services.
[0044] In this context, "participant terminal 120" can be understood as an application installed on a computer device, or it can be understood as a hardware device corresponding to a server. More specifically, the hardware device corresponding to a server refers to computer devices such as smartphones, interactive whiteboards, and personal computers. It should be understood that when participant terminal 120 is a smartphone, interactive whiteboard, or other hardware device, users can install matching applications on participant terminal 120, and can also access web-based applications on participant terminal 120.
[0045] In this embodiment, each participant terminal 120 is equipped with an audio conferencing application. Participants can join the same meeting room, such as a remote meeting, by clicking to access the audio conferencing application installed on their respective terminals 120. The aforementioned meeting room refers to a chat room implemented using internet technology and a server, typically equipped with audio and video playback control functions. Each participant can speak and discuss within the meeting room through their respective terminal 120.
[0046] Specifically, participants speaking in the conference room collect their own audio data in real time through their respective participant terminals 120, and then send the audio data to the server 110; the server 110 forwards the audio data to other participant terminals 120 for playback, thereby enabling remote conference discussions among all participants.
[0047] It is worth mentioning that, Figure 1 The application scenario described is merely an exemplary one and is not intended to limit the scope of the present invention. The solution of the present invention can also be applied to other forms of audio processing applications, and will be processed using existing data processing methods according to the actual application scenario. For example, in audio-visual conferencing, each participant terminal 120 also collects its own screen data and sends the screen data to the server 110, which then mixes the screen data and forwards it to each participant terminal 120 for display. This will not be described in detail further.
[0048] During audio conferences, it is often necessary to record the participants' statements. Related technologies typically involve recording the audio of the meeting, which users can then access as a file.
[0049] However, due to poor recording quality or unclear audio from the speaker, users cannot accurately understand the meeting content through the audio files, thus affecting their comprehension of the meeting content.
[0050] To solve the above-mentioned technical problems, in another related technology, after the server 110 obtains the audio data of each participant, it also mixes the audio data of each participant. After the meeting ends, it generates a mixed audio data and finally performs speech recognition based on the mixed audio data to obtain the meeting text.
[0051] However, the meeting text obtained by mixing audio data cannot effectively distinguish the content of each participant's speech, which still affects the user's understanding of the meeting content.
[0052] Therefore, embodiments of this application propose audio processing methods, apparatus, storage media, and electronic devices.
[0053] The following will be combined with the appendix Figures 2 to 9 This application provides a detailed description of the audio processing method provided in its embodiments.
[0054] Please see Figure 2 This is a flowchart of an audio processing method provided in one embodiment of this application. The audio processing method provided in this embodiment includes the following steps:
[0055] Step S101: Obtain the audio data of each participant, the participant identifier corresponding to each audio data, the first time identifier of each audio data, the mixed audio data, and the second time identifier of the mixed audio data; the mixed audio data is obtained by mixing the audio data of each participant.
[0056] It is understood that the audio processing method in this application embodiment can be executed by one of the participating terminals, by the server, or even by other electronic devices such as recording devices connected to the server or the participating terminal. This application embodiment uses the server as the execution subject for illustrative purposes.
[0057] The server begins recording the audio conference upon detecting the start of the meeting. Specifically, when the server detects that at least one participant has entered the meeting room or arrived at the scheduled start time, it begins receiving audio data transmitted by the participants through their respective endpoints, obtaining the participant identifier corresponding to the audio data and the first time identifier of each audio data. The server also mixes the audio data from all participants to obtain mixed audio data, and simultaneously determines the second time identifier of the mixed audio data.
[0058] It should be understood that the server can mix the audio data in real time, thereby obtaining the mixed speech text data associated with the participant identifiers in real time, and then displaying the mixed speech text data in the conference room in real time. The server can also mix the acquired audio data after the meeting ends, thereby obtaining the mixed speech text data recording the content of this meeting after the meeting concludes. This application embodiment does not impose any limitations on this.
[0059] The first time identifier is used to identify the audio data. In an optional embodiment, the first time identifier can be a point in time; specifically, the first time identifier can be the time when a participant joins the meeting, the time when a participant turns on their microphone, or even a midpoint in the audio data. The first time identifier can be a time determined based on the server's system time or a time determined based on the system time of the participant's device where the audio data is located.
[0060] It should be understood that the first time marker is the time a participant joins the meeting. Considering that participants may join and then leave the meeting multiple times, an audio record is made for each participant joining the meeting, along with the corresponding first time marker. Similarly, the first time marker is the time a participant turns on their microphone. Considering that participants may turn on and off their microphones multiple times during the meeting, an audio record is made for each time a participant turns on their microphone, along with the corresponding first time marker.
[0061] The second time identifier is used to identify the mixed audio data. In an optional embodiment, the second time identifier can be a point in time; specifically, the second time identifier can be the meeting start time or the time when the first participant enters the meeting room. The second time identifier can be a time determined based on the server's system time or a time determined based on the system time of the participant's terminal that triggered the meeting.
[0062] It should be understood that the audio data and mixed audio data in this application include not only the audio itself, but also the audio time. The audio time includes at least the various times during the playback of this audio segment. It can be a time identified by duration, or it can be a specific time. For example, if the audio is 10 seconds long, the time identified by the 10-second duration can be displayed from the beginning to the end of the audio playback, such as the 1st second, the 2nd second, ... the 10th second; or, the specific time can be displayed from the beginning to the end of the audio playback, such as January 1, 1:01, 1:02, ... January 1, 1:10.
[0063] Step S102: Identify each audio data point to obtain the spoken text data.
[0064] In an optional implementation, the server stores a speech recognition algorithm, and the server uses the speech recognition algorithm to recognize the audio data of each participant to obtain the speech text data of each participant.
[0065] In another optional implementation, the server stores a trained speech recognition model. The server inputs the audio data of each participant into the speech recognition model to obtain the spoken text data of each participant. The trained speech recognition model can be obtained by training an existing neural network model on an existing speech dataset.
[0066] In an optional embodiment, in order to accelerate the efficiency of acquiring speech text data and reduce the processing burden of a single server, at least two servers are used for processing, one of which is a business processing server and the other is an audio data processing server. The business server receives the audio data of each participant in real time, mixes the audio data of each participant, and forwards the audio data to the audio data processing server. The audio data processing server identifies the audio data of each participant to obtain the speech text data of each participant.
[0067] It is understandable that, since audio data includes the audio itself as well as the audio duration, the resulting spoken text data will also be the text corresponding to the audio duration.
[0068] Step S103: Based on each first time identifier and second time identifier, align each participant identifier, speech text data and mixed audio data to obtain the mixed speech text data associated with the participant identifier.
[0069] Understandably, since the mixed audio data is obtained by mixing the audio data of each participant, the position of each audio data in the mixed audio data can be determined based on the first and second time markers. This allows the participant identifiers and speech text data corresponding to each audio data to be aligned with the mixed audio data. The aligned participant identifiers and speech text data are then spliced together to obtain the mixed speech text data associated with the participant identifiers.
[0070] This application embodiment acquires audio data of each participant, participant identifiers corresponding to each audio data, first time identifiers of each audio data, mixed audio data, and second time identifiers of the mixed audio data; the mixed audio data is obtained by mixing the audio data of each participant; each audio data is identified to obtain speech text data; based on each first time identifier and second time identifier, each participant identifier, speech text data, and mixed audio data are aligned to obtain mixed speech text data associated with the participant identifier, thereby effectively distinguishing the speech content of each participant based on the mixed speech text data, making it easier for users to understand the meeting content.
[0071] Please see Figure 3 In an optional embodiment, step S101, which involves obtaining the audio data of each participant and the participant identifier corresponding to each audio data, includes:
[0072] Step S111: Receive speaking operation requests from each participant; the speaking operation request includes the participant's identifier.
[0073] In an optional embodiment, the speaking operation request can be a request for each participant to join the meeting. Specifically, after a participant joins the meeting based on the meeting identifier, the participant terminal will obtain the pre-stored participant identifier or receive the participant identifier input by the participant, and generate a speaking operation request based on the participant identifier and send it to the server.
[0074] In another optional embodiment, the speaking operation request can also be a microphone activation request from each participant. Specifically, after a participant triggers the microphone activation control, the participant will obtain the pre-stored participant identifier and generate a speaking operation request based on the participant identifier, which will then be sent to the server.
[0075] Step S112: Based on the participant identifier of each participant, the data transmission channel where each speaking operation request is located, and the audio data transmitted by each data transmission channel, determine the audio data of each participant and the participant identifier corresponding to each audio data.
[0076] In an optional embodiment, each participant joining the conference will establish a data transmission channel with the server. The server can identify the data transmission channel established with each participant to determine the corresponding speaking operation request. Then, based on the participant identifier carried in the speaking operation request and the audio data transmitted through the data transmission channel, the server can determine the audio data of each participant and the participant identifier corresponding to each audio data.
[0077] In another optional embodiment, each participant joining the conference will establish a data transmission channel with the server. Each time a participant transmits audio data, it will carry its participant identifier. The data transmission channel of the audio data can be determined based on the participant identifier, thereby determining the corresponding speaking operation request. Then, based on the participant identifier carried in the speaking operation request and the audio data transmitted through the data transmission channel, the audio data of each participant and the participant identifier corresponding to each audio data can be determined.
[0078] Based on the participant identifier of each participant, the data transmission channel where each speaking operation request is located, and the audio data transmitted through each data transmission channel, the embodiments of this application can conveniently and accurately obtain the audio data of each participant and the participant identifier corresponding to each audio data.
[0079] Please see Figure 4 In an optional embodiment, step S101, which involves obtaining the first time identifier of each audio data point, includes:
[0080] Step S121: Receive speaking requests from each participant.
[0081] Step S122: The time when each participant's speaking operation request is received is determined as the first time identifier of each audio data transmitted based on the speaking operation request.
[0082] In an optional embodiment, the speaking operation request may be a request from each participant to join the meeting, and the first time identifier of each audio data is the time when the server receives the request from each participant to join the meeting.
[0083] In another optional embodiment, the speaking operation request can also be a request from each participant to turn on their microphone, in which case the first time of each audio data is marked as the time when the server receives the request from each participant to turn on their microphone.
[0084] This application embodiment takes into account network latency and differences in system time between different devices, and uniformly sets the first time identifier of each audio data to be determined by the server system time. This can avoid time errors caused by device differences and network latency, thereby ensuring that the participant identifier, speech text data and mixed audio data of each subsequent audio data are aligned.
[0085] Please see Figure 5 In an optional embodiment, step S101, which involves obtaining the second time identifier of the mixed audio data, includes:
[0086] Step S131: Listen to the start of the meeting.
[0087] Step S132: Determine the time corresponding to the start of the meeting as the second time identifier of the mixed audio data.
[0088] Understandably, the meeting start time can be either the scheduled start time or the time when the first participant joins the meeting room.
[0089] In this embodiment, the server determines the time corresponding to the start of the meeting as the second time identifier of the mixed audio data, and then sets the second time identifier of the mixed audio data to be determined by the server system time, which can ensure that the participant identifiers, speech text data and mixed audio data of subsequent audio data are aligned with the mixed audio data.
[0090] Please see Figure 6 In an optional embodiment, step S103, which involves aligning each participant identifier, speech text data, and mixed audio data according to each first time identifier and second time identifier to obtain mixed speech text data associated with the participant identifier, includes:
[0091] Step S1031: Obtain the time offset of each audio data relative to the mixed audio data based on each first time identifier and the second time identifier.
[0092] Step S1032: Based on each time offset, align the participant identifier and corresponding speech text data corresponding to each audio data with the mixed audio data to obtain the mixed speech text data associated with the participant identifier.
[0093] It is understandable that, since the mixed audio data is obtained by mixing various audio data, and each audio data and the mixed audio data itself carry time data, after obtaining the time offset of each audio data relative to the mixed audio data based on each first time identifier and second time identifier, the individual audio data and the mixed audio data can be aligned with each other based on time. Since each audio data can be identified as speech text data, and each audio data corresponds to a participant identifier, the participant identifier and the corresponding speech text data of each audio data can be aligned with the mixed audio data to obtain the mixed speech text data associated with the participant identifier.
[0094] For example, such as Figure 7 As shown, when the first time identifier is the microphone activation time of each participant and the second time identifier is the meeting start time, the time offset of each audio data relative to the mixed audio data is obtained based on the first time identifier and the second time identifier. Then, the audio data of each participant A, participant B, participant C and participant D can be aligned with the mixed audio data based on time. This aligns the participant identifiers and corresponding speech text data of participants A, B, C and D with the mixed audio data, and obtains the mixed speech text data associated with the participant identifier.
[0095] It is understandable that, such as Figure 7 As shown, considering the recording time and size of the mixed audio data, the mixed audio will be segmented at intervals to obtain several mixed audio segments. The various mixed audio segments will be spliced together before the alignment operation is performed.
[0096] This application aligns each audio data point with the mixed audio data based on the time offset of each audio data point relative to the mixed audio data. The corresponding participant identifier and corresponding speech text data are also aligned with the mixed audio data. Compared to the solution of directly identifying the mixed audio data as mixed speech text data, this application can avoid the problem that the speech content of each participant cannot be effectively distinguished based on the mixed audio data when all participants speak at the same time. Compared to the solution of directly splicing the speech text data of each audio data point into mixed speech text data, this application does not need to find the splicing time point and can avoid the problem of mixed speech text errors caused by splicing time point errors.
[0097] In an optional embodiment, step S1032, which involves obtaining the time offset of each audio data relative to the mixed audio data based on each first time identifier and a second time identifier, includes:
[0098] Step S1321: Subtract the second time identifier from each first time identifier to obtain the time offset of each audio data relative to the mixed audio data.
[0099] In this embodiment of the application, the time offset of each audio data relative to the mixed audio data can be conveniently obtained by subtracting the second time identifier from each first time identifier.
[0100] Please see Figure 8 In an optional embodiment, the audio processing method of this application further includes:
[0101] Step S104: When playing the mixed audio data, process the mixed speech text data associated with the participant identifier into subtitles with the participant identifier, and display the subtitles with the participant identifier.
[0102] Specifically, the mixed-stream speech text data is displayed as subtitles in the video stream data, and corresponding participant identifiers are added to each subtitle to distinguish the subtitles corresponding to the audio data of different participants. Specifically, when playing the mixed-stream audio data of the audio conference, a subtitle interface can be displayed, on which the corresponding subtitles are played. When video data is included, such as... Figure 9 As shown, participant logos and subtitles can be displayed below the video stream data and / or on the right side of the video stream data, and this embodiment of the invention does not impose any limitations on this.
[0103] In this embodiment of the application, when playing mixed audio data, the mixed speech text data associated with the participant identifier is processed into subtitles with the participant identifier added, and the subtitles with the participant identifier added are displayed. This allows users to accurately obtain the meeting content by combining the mixed audio data and the subtitles.
[0104] This application also provides an audio processing system, including: a server and a plurality of participating terminals; each participating terminal is connected to the server;
[0105] Each participant sends its own audio data and the participant identifier corresponding to each audio data to the server.
[0106] The server is used to obtain the audio data of each participant, the participant identifier corresponding to each audio data, the first time identifier of each audio data, the mixed audio data, and the second time identifier of the mixed audio data; the mixed audio data is obtained by mixing the audio data of each participant.
[0107] The server is used to identify each audio data point and obtain the spoken text data;
[0108] The server is used to align the participant identifiers, speech text data, and mixed audio data according to the first and second time identifiers to obtain the mixed speech text data associated with the participant identifiers.
[0109] The audio processing system and the audio processing method provided in this application belong to the same concept, and their implementation process can be found in the method embodiment, which will not be repeated here.
[0110] Please see Figure 10 This is a schematic diagram of an audio processing apparatus provided in one embodiment of this application. The apparatus 200 includes:
[0111] The data acquisition module 201 is used to acquire the audio data of each participant, the participant identifier corresponding to each audio data, the first time identifier of each audio data, the mixed audio data, and the second time identifier of the mixed audio data; the mixed audio data is obtained by mixing the audio data of each participant.
[0112] The text recognition module 202 is used to recognize each audio data and obtain the spoken text data;
[0113] The speech text acquisition module 203 is used to align the participant identifiers, speech text data and mixed audio data according to the first time identifier and the second time identifier, and obtain the mixed speech text data associated with the participant identifiers.
[0114] It should be noted that the audio processing apparatus provided in this application embodiment is only illustrated by the above-described division of functional modules when executing the audio processing method. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the audio processing apparatus provided in this application embodiment and the audio processing method provided in this application embodiment belong to the same concept, and the implementation process is detailed in the method embodiment, which will not be repeated here.
[0115] The audio processing apparatus embodiments provided in this application can be applied to computer devices, such as servers. These apparatus embodiments can be implemented through software, hardware, or a combination of both. Taking software implementation as an example, as a logical device, it is formed by a processor that reads and executes corresponding computer program instructions from memory. From a hardware perspective, the computer device in which it resides may include a processor and memory, which are interconnected via a data bus or other known methods.
[0116] Please see Figure 11 This is a schematic diagram of the structure of an electronic device provided in one embodiment of this application. Figure 11 As shown, the electronic device 300 can specifically be a computer, mobile phone, tablet computer, interactive flat panel, etc. In the exemplary embodiment of this application, the electronic device 300 is a server. The electronic device 300 may include: at least one processor 310, at least one memory 320, at least one network interface 330, user interface 340, and at least one communication bus 350.
[0117] The communication bus 350 is used to enable communication between these components.
[0118] The user interface 340 may include a display screen and a camera; the user interface 340 may also include standard wired and wireless interfaces.
[0119] The network interface 330 may optionally include a standard wired interface and a wireless interface (such as a Wi-Fi interface).
[0120] The processor 310 may include one or more processing cores. The processor 310 connects to various parts within the electronic device 300 using various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 320, and by calling data stored in the memory 320. Optionally, the processor 310 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 310 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also be implemented as a separate chip without being integrated into the processor 310.
[0121] The memory 320 may include random access memory (RAM) or read-only memory. Optionally, the memory 320 may include a non-transitory computer-readable storage medium. The memory 320 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 320 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 320 may also be at least one storage device located remotely from the aforementioned processor 310. Figure 11 As shown, the memory 320, which serves as a computer storage medium, may include an operating system, a network communication module, and a user.
[0122] exist Figure 11 In the electronic device 300 shown, the user interface 340 is mainly used to provide an input interface for the user and to obtain the user input data; while the processor 310 can be used to call the server operation application stored in the memory 320, such as an audio processing program; and execute the relevant operations of any audio processing method in the above embodiments, with corresponding functions and beneficial effects.
[0123] This application also provides a computer-readable storage medium storing a computer program, the instructions of which are adapted to be loaded by a processor and executed by the audio processing method steps described above. For details of the execution process, please refer to the specific descriptions in the embodiments, which will not be repeated here. The device containing the storage medium can be an electronic device such as a personal computer, laptop computer, smartphone, or tablet computer.
[0124] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative, wherein the components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0125] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0126] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function selected in one or more boxes.
[0127] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function selected in one or more boxes.
[0128] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0129] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0130] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0131] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0132] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. An audio processing method, characterized in that, Includes the following steps: The system acquires audio data from each participant, participant identifiers corresponding to each audio data, first time identifiers for each audio data, mixed audio data, and second time identifiers for the mixed audio data; the mixed audio data is obtained by mixing the audio data from each participant. Identify each of the audio data points to obtain the spoken text data; Based on each of the first time identifiers and the second time identifiers, the time offset of each audio data relative to the mixed audio data is obtained; Based on the time offsets, the participant identifiers and corresponding speech text data corresponding to each audio data are aligned with the mixed audio data. Then, the aligned participant identifiers and speech text data are spliced together to obtain the mixed speech text data associated with the participant identifiers.
2. The audio processing method according to claim 1, characterized in that: The step of obtaining the time offset of each audio data relative to the mixed audio data based on each of the first time identifiers and the second time identifiers includes: Subtract the second time identifier from each of the first time identifiers to obtain the time offset of each audio data relative to the mixed audio data.
3. The audio processing method according to any one of claims 1 to 2, characterized in that... , The step of obtaining the first time identifier of each of the audio data includes: Receive speaking requests from each participant; The time when each participant's speaking request is received is determined as the first time identifier of each audio data transmitted based on the speaking request.
4. The audio processing method according to any one of claims 1 to 2, characterized in that, The step of obtaining the second time identifier of the mixed audio data includes: Start monitoring the meeting; The time corresponding to the start of the meeting is determined as the second time identifier of the mixed audio data.
5. The audio processing method according to claim 1, characterized in that, The steps of obtaining the audio data of each participant and the participant identifier corresponding to each audio data include: Receive speaking requests from each participant; the speaking request includes the participant's identifier; Based on the participant identifier of each participant, the data transmission channel where each speaking operation request is located, and the audio data transmitted through each data transmission channel, determine the audio data of each participant and the participant identifier corresponding to each audio data.
6. The audio processing method according to claim 1, characterized in that, It also includes the following steps: When playing the mixed audio data, the mixed speech text data associated with the participant identifier is processed into subtitles with the participant identifier added, and the subtitles with the participant identifier added are displayed.
7. An audio processing device, characterized in that, include: The data acquisition module is used to acquire the audio data of each participant, the participant identifier corresponding to each audio data, the first time identifier of each audio data, the mixed audio data, and the second time identifier of the mixed audio data; the mixed audio data is obtained by mixing the audio data of each participant. The text recognition module is used to recognize each audio data item and obtain the spoken text data; The speech text acquisition module is used to obtain the time offset of each audio data relative to the mixed audio data based on each first time identifier and the second time identifier; based on each time offset, align the participant identifier and the corresponding speech text data of each audio data with the mixed audio data; and then concatenate and integrate the aligned participant identifier and speech text data to obtain the mixed speech text data associated with the participant identifier.
8. An audio processing system, comprising: The server and several participating terminals; each participating terminal is connected to the server; characterized in that: Each participating terminal sends its own audio data and the participant identifier corresponding to each audio data to the server; The server is used to acquire audio data from each participant, participant identifiers corresponding to each audio data, first time identifiers of each audio data, mixed audio data, and second time identifiers of the mixed audio data; the mixed audio data is obtained by mixing the audio data from each participant. The server is used to identify each audio data item and obtain spoken text data; The server is used to obtain the time offset of each audio data relative to the mixed audio data according to each first time identifier and the second time identifier; according to each time offset, the participant identifier and the corresponding speech text data corresponding to each audio data are aligned with the mixed audio data; and the aligned participant identifier and speech text data are then spliced together to obtain the mixed speech text data associated with the participant identifier.
9. An electronic device comprising a processor and a memory; characterized in that, The memory stores a computer program adapted to be loaded by the processor and executed as described in any one of claims 1 to 6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the audio processing method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Conference sub-role speech synthesis method and device, computer equipment and storage medium
CN110322869A
Voice processing method and device
CN111835529A