Audio processing method, mobile terminal, system and storage medium

The mobile terminal has a built-in audio acquisition module and processing application, which automatically collects and processes audio data to generate meeting records, solving the problem of convenient recording and realizing efficient and flexible meeting record generation.

CN115132202BActive Publication Date: 2025-10-03BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110328410.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-26
Publication Date
2025-10-03
Estimated Expiration
2041-03-26

AI Technical Summary

Technical Problem

Existing meeting recording methods are not convenient enough. Notes cannot fully restore the meeting process. Recording and video equipment need to be prepared in advance, which consumes time and effort.

Method used

The mobile terminal has a built-in audio acquisition module, voice interaction application and text processing application. By turning on the audio processing function, it automatically collects, processes and outputs audio data to generate meeting records.

Benefits of technology

Efficient meeting records can be easily obtained without additional equipment, reducing manual processing time. It supports voice recognition, translation and voiceprint partition display, improving recording accuracy and flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115132202B_ABST
    Figure CN115132202B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an audio processing method, a mobile terminal, a system and a storage medium. The audio processing method is applied to a mobile terminal, and the mobile terminal includes a voice interaction application and a text processing application. The audio processing method includes: when detecting an activation instruction for the audio processing function of the mobile terminal, obtaining first audio data based on the audio acquisition module of the mobile terminal; processing the first audio data through the voice interaction application of the mobile terminal to obtain an audio processing result; wherein, the voice interaction application corresponds to at least one type of audio processing module, and different types of audio processing modules are used to perform different processing on the first audio data; starting the text processing application of the mobile terminal, and outputting the audio processing result through the text processing application. In this way, even if a meeting is called at short notice, there is no need to temporarily prepare additional recording and video equipment, which provides convenience for the formation of meeting minutes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of audio technology, and in particular to an audio processing method, a mobile terminal, a system, and a storage medium. Background Art

[0002] A meeting is an organized, directed, and purposeful deliberative activity, conducted at a defined time and place according to a set procedure. Meetings are a common social phenomenon, occurring in almost every organized place. Meeting minutes are particularly important. Meeting minutes are the recorder's record of the meeting's organization and specific content, forming a record. "Minutes" can be divided into detailed and brief. Brief minutes capture the main points of the meeting, including important or major speeches; detailed minutes require that all recorded items be complete, with speeches recorded in detail and completeness. Recordings that include these contents must be recorded in writing, audio, or video. Audio and video recordings are usually just a means to a successful meeting, ultimately requiring the recorded content to be translated into text. Written minutes often rely on audio and video recordings to ensure that the recorded content best captures the meeting context.

[0003] However, recording meetings is often inconvenient. The most basic method relies on handwritten notes, which, due to hand speed limitations, can only provide a cursory record and cannot fully reconstruct the entire meeting process. Currently, the most common method is transcribed recordings. After the meeting, the recorder reviews the recordings and manually converts the meeting content into text for easy storage and access. This method not only requires specialized recording and video equipment in advance, but also consumes a significant amount of time and effort on the recorder's part. Summary of the Invention

[0004] The present disclosure provides an audio processing method, a mobile terminal, a system and a storage medium.

[0005] According to a first aspect of an embodiment of the present disclosure, there is provided an audio processing method, which is applied to a mobile terminal, wherein the mobile terminal includes a voice interaction application and a text processing application. The method includes:

[0006] When an activation instruction for the audio processing function of the mobile terminal is detected, acquiring first audio data based on an audio acquisition module of the mobile terminal;

[0007] Processing the first audio data by the voice interaction application of the mobile terminal to obtain an audio processing result; wherein the voice interaction application corresponds to at least one type of audio processing module, and the different types of audio processing modules are respectively used to perform different processing on the first audio data;

[0008] The text processing application of the mobile terminal is started, and the audio processing result is output through the text processing application.

[0009] Optionally, outputting the audio processing result through the text processing application includes:

[0010] The audio processing result is output in a text format by the text processing application.

[0011] Optionally, the at least one type of audio processing module includes: a speech recognition module; the processing of the first audio data by the voice interaction application of the mobile terminal to obtain an audio processing result includes:

[0012] The first audio data is recognized by the voice recognition module of the voice interaction application to obtain semantic content corresponding to the first audio data.

[0013] Optionally, the at least one type of audio processing module further includes: a translation module; and the processing of the first audio data by the voice interaction application of the mobile terminal to obtain an audio processing result further includes:

[0014] When it is determined that the translation function of the mobile terminal is turned on, the semantic content is translated by the translation module of the voice interaction application to obtain translation content expressed in a target language.

[0015] Optionally, the at least one type of audio processing module further includes: a voiceprint processing module; and the method further includes:

[0016] Inputting the first audio data into the voiceprint processing module of the voice interaction application to obtain the voiceprint features of each sound-making object;

[0017] Determining feature identifiers of various parts of the audio processing result based on the voiceprint features of each of the sound-making objects;

[0018] The text processing application displays the audio processing results in partitions on the current interface according to the feature identifiers.

[0019] Optionally, the mobile terminal includes: a plurality of the audio acquisition modules, and the audio acquisition surfaces of the audio acquisition modules face different directions;

[0020] The acquiring of first audio data by the audio acquisition module based on the mobile terminal includes:

[0021] Based on the multiple audio acquisition modules, multiple groups of second audio data are collected, and the multiple groups of second audio data are fused to obtain the first audio data.

[0022] Optionally, the method further includes:

[0023] In response to an editing instruction for the audio processing result on the text processing application, the audio processing result is edited.

[0024] Optionally, the method further includes:

[0025] Sending the authentication information of the mobile terminal to the middle platform device through the communication module, so that the middle platform device authenticates the mobile terminal;

[0026] Wherein, the authentication information at least includes: the device identification of the mobile terminal.

[0027] Optionally, the processing the first audio data by the voice interaction application of the mobile terminal to obtain an audio processing result includes:

[0028] When it is determined that the mobile terminal authentication is passed, processing the first audio data by a voice interaction application of the mobile terminal to obtain the audio processing result;

[0029] The mobile terminal authentication includes: a preset identifier in a preset identifier list of the middle station device has a preset identifier that matches the device identifier.

[0030] Optionally, at least one type of audio processing module corresponding to the voice interaction application is located in the middle platform device.

[0031] According to a second aspect of an embodiment of the present disclosure, there is provided an audio processing method, which is applied to an audio processing system, the audio processing system including a mobile terminal and a mid-stage device, the method comprising:

[0032] When the mobile terminal detects an activation instruction for the audio processing function, the mobile terminal acquires the first audio data based on the audio acquisition module of the mobile terminal;

[0033] The mobile terminal obtains, from the middle platform device, an audio processing result obtained by processing the first audio data through a voice interaction application; wherein the voice interaction application corresponds to at least one type of audio processing module, the audio processing module is located in the middle platform device, and the different types of audio processing modules are respectively used to perform different processing on the first audio data;

[0034] The mobile terminal starts a text processing application and outputs the audio processing result through the text processing application.

[0035] Optionally, the mobile terminal obtains, from the middleware device through a voice interaction application, an audio processing result obtained by processing the first audio data, including:

[0036] The mobile terminal sends authentication information to the middle station device through the communication module; wherein the authentication information includes: the device identification of the mobile terminal;

[0037] The middle station device compares the device identifier with the preset identifiers in the preset identifier list to determine whether there is a preset identifier in the preset identifier list that matches the device identifier;

[0038] When the middle station device determines that there is a preset identifier matching the device identifier in the preset identifier list, it determines that the authentication of the mobile terminal is successful.

[0039] According to a third aspect of an embodiment of the present disclosure, there is provided a mobile terminal, including:

[0040] an audio acquisition module, configured to acquire first audio data when detecting an activation instruction for an audio processing function of the mobile terminal;

[0041] a voice interaction application configured to process the first audio data to obtain an audio processing result; wherein the voice interaction application corresponds to at least one type of audio processing module, and the different types of audio processing modules are respectively used to perform different processing on the first audio data;

[0042] The text processing application is configured to output the audio processing result.

[0043] Optionally, the text processing application is further configured to:

[0044] The audio processing result is output in text format.

[0045] Optionally, the at least one type of audio processing module includes: a speech recognition module; the speech recognition module is configured to:

[0046] The first audio data is recognized by the voice recognition module of the voice interaction application to obtain semantic content corresponding to the first audio data.

[0047] Optionally, the at least one type of audio processing module further includes: a translation module; the translation module is configured to:

[0048] When it is determined that the translation function of the mobile terminal is turned on, the semantic content is translated by the translation module of the voice interaction application to obtain translation content expressed in a target language.

[0049] Optionally, the at least one type of audio processing module further includes: a voiceprint processing module; the voiceprint processing module is configured to:

[0050] Obtain the voiceprint characteristics of each sound-making object;

[0051] Determining feature identifiers of various parts of the audio processing result based on the voiceprint features of each of the sound-making objects;

[0052] The text processing application is further configured to display the audio processing results in partitions on the current interface according to the feature identifiers.

[0053] Optionally, the mobile terminal includes: a plurality of the audio acquisition modules, and the audio acquisition surfaces of the audio acquisition modules face different directions;

[0054] The plurality of audio acquisition modules are further configured as follows:

[0055] Collect multiple sets of second audio data, and perform fusion processing on the multiple sets of second audio data to obtain the first audio data.

[0056] Optionally, the text processing application is further configured to:

[0057] In response to an editing instruction for the audio processing result, the audio processing result is edited.

[0058] Optionally, the mobile terminal further includes:

[0059] a communication module configured to send the authentication information of the mobile terminal to the middle platform device so that the middle platform device authenticates the mobile terminal;

[0060] Wherein, the authentication information at least includes: the device identification of the mobile terminal.

[0061] Optionally, the voice interaction application is further configured to:

[0062] When it is determined that the mobile terminal authentication is passed, processing the first audio data to obtain the audio processing result;

[0063] The mobile terminal authentication includes: a preset identifier in a preset identifier list of the middle station device has a preset identifier that matches the device identifier.

[0064] Optionally, at least one type of audio processing module corresponding to the voice interaction application is located in the middle platform device.

[0065] According to a fourth aspect of an embodiment of the present disclosure, there is provided an audio processing system, including:

[0066] Mobile terminals and middleware devices;

[0067] The mobile terminal is configured to acquire first audio data based on an audio acquisition module of the mobile terminal when an activation instruction for an audio processing function is detected;

[0068] Obtaining, from the middle platform device, an audio processing result obtained by processing the first audio data through a voice interaction application; wherein the voice interaction application corresponds to at least one type of audio processing module, the audio processing module is located in the middle platform device, and different types of audio processing modules are respectively used to perform different processing on the first audio data;

[0069] The mobile terminal is further configured to start a text processing application and output the audio processing result through the text processing application.

[0070] Optionally, the mobile terminal is further configured to send authentication information to the middle station device through the communication module; wherein the authentication information includes: a device identifier of the mobile terminal;

[0071] The middle station device is configured to compare the device identifier with a preset identifier in a preset identifier list to determine whether there is a preset identifier in the preset identifier list that matches the device identifier;

[0072] When it is determined that there is a preset identifier in the preset identifier list that matches the device identifier, it is determined that the mobile terminal authentication is successful.

[0073] According to a fifth aspect of an embodiment of the present disclosure, there is provided a mobile terminal, including:

[0074] processor;

[0075] a memory configured to store processor-executable instructions;

[0076] The processor is configured to implement the steps of any one of the audio processing methods in the first aspect or the second aspect when executing.

[0077] According to the sixth aspect of an embodiment of the present disclosure, a non-temporary computer-readable storage medium is provided. When the instructions in the storage medium are executed by a processor of a mobile terminal, the mobile terminal is enabled to execute any one of the audio processing methods in the first aspect or the second aspect mentioned above.

[0078] The technical solutions provided by the embodiments of the present disclosure may have the following beneficial effects:

[0079] In an embodiment of the present disclosure, a mobile terminal can be provided. Since the mobile terminal itself includes an audio acquisition module, a voice interaction application and a text processing application, when the mobile terminal of the present disclosure detects an instruction to start the audio processing function, it can obtain first audio data based on the mobile terminal's own audio acquisition module, and process the first audio data through its own voice interaction application to obtain an audio processing result. After obtaining the audio processing result, it can start its own text processing application and output the audio processing result through the text processing application.

[0080] Since the mobile terminal in the present disclosure can install applications with various functions and has communication functions, it can be used by users in daily life. Users generally carry it with them. In this way, even if a meeting is notified at short notice, as long as the user carries the mobile terminal, he or she only needs to turn on the audio processing function of the mobile terminal during the meeting, and can call the mobile terminal's own audio acquisition module to perform audio acquisition, and process the collected first audio data through the voice interaction application to obtain an audio processing result, and finally output the audio processing result through a text processing application, without the need to temporarily prepare additional recording and video equipment, which provides convenience for the formation of meeting minutes.

[0081] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0083] Figure 1 The figure is a flowchart of an audio processing method according to an exemplary embodiment.

[0084] Figure 2 The figure is a schematic diagram showing the architecture of a mobile terminal according to an exemplary embodiment.

[0085] Figure 3 The figure is a flowchart of another audio processing method according to an exemplary embodiment.

[0086] Figure 4 The figure is a schematic diagram showing the architecture of an audio processing system according to an exemplary embodiment.

[0087] Figure 5 The figure is a structural block diagram of a mobile terminal according to an exemplary embodiment.

[0088] Figure 6 The figure is a block diagram of an audio processing system according to an exemplary embodiment.

[0089] Figure 7 is a block diagram showing a mobile terminal 1200 according to an exemplary embodiment.

[0090] Figure 8 FIG. 1 is a block diagram of a mid-station device 1300 according to an exemplary embodiment. DETAILED DESCRIPTION

[0091] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.

[0092] Figure 1 is a flowchart of an audio processing method according to an exemplary embodiment. Figure 1 As shown, the audio processing method is applied to a mobile terminal, which includes a voice interaction application and a text processing application. The method mainly includes the following steps:

[0093] In step 101, when an activation instruction for the audio processing function of the mobile terminal is detected, first audio data is acquired based on an audio acquisition module of the mobile terminal;

[0094] In step 102, the first audio data is processed by the voice interaction application of the mobile terminal to obtain an audio processing result; wherein the voice interaction application corresponds to at least one type of audio processing module, and the different types of audio processing modules are used to perform different processing on the first audio data;

[0095] In step 103, the text processing application of the mobile terminal is started, and the audio processing result is output through the text processing application.

[0096] Here, the mobile terminal can be a terminal device equipped with an audio acquisition module, a voice interaction application, and a text processing application. The mobile terminal can include at least one of the following: a mobile phone, a tablet computer, a laptop computer, etc. The audio processing function can be used to capture audio data from participants during a meeting or conversation and generate corresponding meeting records based on the audio data.

[0097] Figure 2 FIG. 1 is a schematic diagram showing an architecture of a mobile terminal according to an exemplary embodiment. Figure 2As shown, the mobile terminal 200 includes: an audio processing module 201, a voice interaction application 202 and a text processing application 203, wherein the audio processing module is connected to the interface of the voice interaction application, and a first connection path is established between the audio processing module and the voice interaction application, and the interface of the voice interaction application is connected to the interface of the text processing application, and a second connection path is established between the voice interaction application and the text processing application.

[0098] During the implementation process, the audio processing module can transmit the acquired first audio data to the voice interaction application through the first connection path. After receiving the first audio data, the voice interaction application processes the first audio data through the audio processing module to obtain an audio processing result, and transmits the audio processing result to the text processing application through the second connection path. After receiving the audio processing result, the text processing application outputs the audio processing result according to the set format.

[0099] In other embodiments, the mobile terminal may further include a communication module, which is used to communicate with other devices. For example, the communication module may include: a Wireless Fidelity (Wi-Fi) module or a ZigBee module.

[0100] Taking the use of the mobile terminal in a conference scenario as an example, during the meeting, if the mobile terminal detects an instruction to turn on the audio processing function, it can collect audio data from each sound-emitting object based on the audio collection module, and send the collected audio collection data to the voice interaction application. The voice interaction application can process the collected audio data to obtain the audio processing result, and send the audio processing result to the text processing application, and output the audio processing result in a set format through the text processing application. In this way, the corresponding meeting record can be obtained based on the audio data of the participating members. Among them, the sound-emitting object can be one or more, and there is no specific limitation here.

[0101] It should be noted that the audio acquisition module mainly collects audio data. In some embodiments, the audio acquisition module may include: a voice acquisition module and a video acquisition module. It should be noted that the audio acquisition module in the present disclosure may not only be used to support the audio processing function of the mobile terminal, but in other embodiments, the audio acquisition module may also be used alone, for example, it may be used for the communication function and the recording and video recording function of the mobile terminal. Here, the voice acquisition module and the video acquisition module may be modules integrated into one functional module, or they may be two independent modules, which are not specifically limited here. For example, the audio acquisition module may be a module formed by the cooperation of a microphone and a camera.

[0102] The voice interaction application can be a system application for voice interaction that comes with the mobile terminal's operating system, or a third-party application for voice interaction installed on the mobile terminal. As long as it is an application that can perform human-computer interaction through voice, there is no specific limitation here. In other embodiments, the voice interaction application can be used alone. For example, the voice interaction application may include: voice assistant, etc.

[0103] The text processing application can be a system application for text processing that comes with the mobile terminal's operating system, or it can be a third-party application installed on the mobile terminal for text processing. As long as it can perform text processing, there is no specific limitation here. In other embodiments, the text processing application can also be used independently. For example, the text processing application may include: a note application, a notepad application, a Word document application, etc.

[0104] In some embodiments, the instruction to activate the audio processing function of the mobile terminal may be input through a physical button of the mobile terminal, for example, by double-clicking the power button of the mobile terminal to input the instruction to activate, or by long-pressing the physical volume-down button of the mobile terminal to input the instruction to activate, etc. In other embodiments, the instruction to activate the audio processing function of the mobile terminal may also be input through a virtual button of the mobile terminal, which is not specifically limited here.

[0105] In other embodiments, a voice interaction application may be used to input a command to enable the audio processing function of the mobile terminal. For example, after waking up the voice interaction application of the mobile terminal, the command to enable the audio processing function may be input through the voice interaction application after waking up. The command to enable the audio processing function may be a voice command. For example, the command to enable the audio processing function may be "Xiao Ai, please enable the audio processing function." Another example may be "Xiao Ai, please enter the meeting recording mode." Of course, the command to enable the audio processing function of the mobile terminal may also be in other forms, and the user may customize the command as needed, as long as the audio processing function of the mobile terminal can be enabled.

[0106] For another example, a command to turn off the audio processing function can be input through the voice interaction application after waking up. For example, the command to turn off the audio processing function can be "Xiao Ai, please turn off the audio processing function", or for another example, the command to turn off the audio processing function can be "Xiao Ai, please end the meeting recording mode".

[0107] In other embodiments, a command to enable the audio processing function of the mobile terminal may be input through a text processing application. For example, after the text processing application of the mobile terminal is activated, if a preset text is detected in the text processing application, it is determined that a command to enable the audio processing function has been detected. The preset text may include a text containing a word such as "meeting," and the user may customize the text as needed, as long as the audio processing function of the mobile terminal can be enabled.

[0108] In other embodiments, an audio processing control for turning on or off the audio processing function may be added to the operating system of the mobile terminal. For example, an audio processing control for turning on or off the audio processing function may be added to the settings interface of the mobile terminal. In this way, when the user needs to use the audio processing function, he or she can enter the settings interface of the mobile terminal and turn on the audio processing control through a touch operation.

[0109] In other embodiments, the user can also enable or disable the audio processing function through shortcuts during the meeting. For example, the audio processing function can be enabled or disabled through physical buttons, virtual buttons, or gestures on the mobile terminal, which are not specifically limited here.

[0110] In an embodiment of the present disclosure, a mobile terminal can be provided. Since the mobile terminal itself includes an audio acquisition module, a voice interaction application and a text processing application, when the mobile terminal of the present disclosure detects an instruction to start the audio processing function, it can obtain first audio data based on the mobile terminal's own audio acquisition module, and process the first audio data through its own voice interaction application to obtain an audio processing result. After obtaining the audio processing result, it can start its own text processing application and output the audio processing result through the text processing application.

[0111] Since the mobile terminal in the present disclosure can install applications with various functions and has communication functions, it can be used by users in daily life. Users generally carry it with them. In this way, even if a meeting is notified at short notice, as long as the user carries the mobile terminal, he or she only needs to turn on the audio processing function of the mobile terminal during the meeting, and can call the mobile terminal's own audio acquisition module to perform audio acquisition, and process the collected first audio data through the voice interaction application to obtain an audio processing result, and finally output the audio processing result through a text processing application, without the need to temporarily prepare additional recording and video equipment, which provides convenience for the formation of meeting minutes.

[0112] In some embodiments, outputting the audio processing result through the text processing application includes:

[0113] The audio processing result is output in a text format by the text processing application.

[0114] In some embodiments, the audio processing result may be a result obtained by processing the first audio data by a voice interaction application. The audio processing result may be semantic content in the form of speech. During implementation, the voice interaction application may input the semantic content in the form of speech into a text processing application, which converts the semantic content in the form of speech to obtain an audio processing result in text format, and outputs the audio processing result in text format through the text processing application.

[0115] In other embodiments, the semantic content in voice form can also be converted through a voice interaction application to obtain an audio processing result in text format, and the audio processing result in text format can be transmitted to a text processing application, and the audio processing result in text format can be output through the text processing application.

[0116] In the embodiment of the present disclosure, the audio processing results can be output in text format based on the text processing application. In this way, the user can directly edit the audio processing results through the text processing application. For example, the audio processing results in text format can be modified, added and deleted to make the final generated meeting minutes more accurate and complete.

[0117] In some embodiments, the at least one type of audio processing module includes a speech recognition module; and the processing of the first audio data by the voice interaction application of the mobile terminal to obtain an audio processing result includes:

[0118] The first audio data is recognized by the voice recognition module of the voice interaction application to obtain semantic content corresponding to the first audio data.

[0119] Here, the audio processing module may recognize the first audio data based on Automatic Speech Recognition (ASR) technology to obtain semantic content corresponding to the first audio data.

[0120] In other embodiments, the first audio data may be processed using acoustic echo cancellation (AEC) technology. After the echo cancellation is performed, the echo-cancelled first audio data may be recognized using ASR technology, thereby reducing the probability of recording ambient noise. In other embodiments, the first audio data may be processed using ASR technology and natural language processing (NLP) technology to shield noise that is not relevant to the audio processing function.

[0121] In other embodiments, the user may also select the output method of the audio processing result as needed. For example, after obtaining the audio processing result, the voice interaction application may call the audio output module of the mobile terminal and output the audio processing result through the audio output module to increase the flexibility of the output method of the audio processing result.

[0122] In some embodiments, the at least one type of audio processing module further includes: a translation module; and the processing of the first audio data by the voice interaction application of the mobile terminal to obtain an audio processing result further includes:

[0123] When it is determined that the translation function of the mobile terminal is turned on, the semantic content is translated by the translation module of the voice interaction application to obtain translation content expressed in a target language.

[0124] It should be noted that, when an audio processing control is set for the audio processing function of the mobile terminal, an audio translation control for the translation function can be set accordingly, and the audio translation control can be used as a sub-control of the audio processing control. For example, when the audio processing control is in the on state, the audio translation control can be turned on or off. When the audio processing control is in the off state, the audio translation control cannot be turned on or off.

[0125] In other embodiments, the user can also enable or disable the translation function during the meeting through shortcuts. For example, the translation function can be enabled or disabled through physical buttons, virtual buttons, or gestures on the mobile terminal, which are not specifically limited here.

[0126] Here, the translation module can translate the semantic content corresponding to the first audio data using a machine translation model. The machine translation model can be a pre-trained model, such as a bilingual translation model or a multi-language translation model. For another example, the machine translation model can be trained based on a convolutional neural network, which is not specifically limited here.

[0127] In some embodiments, upon determining that the translation function of the mobile terminal is enabled, if the semantic content obtained is expressed in a first language, the semantic content expressed in the first language can be translated into semantic content expressed in the target voice based on the translation module. The first language can be English, and the target language can be Chinese. In other embodiments, the target language can also be other types of languages, for example, English, Japanese, etc.

[0128] In the disclosed embodiment, after obtaining the semantic content corresponding to the first audio data, the translation module of the voice interaction application can directly translate the voice content to obtain the translation content represented by the target voice. This eliminates the need for users to manually translate or download additional translation applications to translate unfamiliar voices, thereby improving the convenience of using mobile terminals.

[0129] In other embodiments, after obtaining the translated content, the semantic content and translated content corresponding to the first audio data can be output simultaneously in text format through a text processing application. By outputting both languages ​​simultaneously, the user can read the content in comparison, thereby improving the readability of the text content.

[0130] In other embodiments, when the semantic content and the translation content are output simultaneously in text format, they may also be output in segmented intervals, that is, each segment of semantic content corresponds to a segment of translation content, to facilitate comparative reading by the user.

[0131] In some embodiments, the at least one type of audio processing module further includes: a voiceprint processing module; the method further includes:

[0132] Inputting the first audio data into the voiceprint processing module of the voice interaction application to obtain the voiceprint features of each sound-making object;

[0133] Determining feature identifiers of various parts of the audio processing result based on the voiceprint features of each of the sound-making objects;

[0134] The text processing application displays the audio processing results in partitions on the current interface according to the feature identifiers.

[0135] Here, the voiceprint processing module can process the first audio data using a voiceprint extraction network model to obtain voiceprint features of each sound-making object. Here, the voice spectrum can be obtained by performing a Fourier transform on the audio signal, and then the spectrum can be input into the voiceprint extraction network model to extract the voiceprint features.

[0136] For example, the voiceprint extraction network model can be composed of a residual network (RESNET), a pooling layer, and a fully connected layer. The pooling layer can include multiple layers, for example, two layers. The voiceprint extraction network model can be a pre-trained model, and the loss function used in model training can be a cross-entropy loss function.

[0137] In the embodiment of the present disclosure, after obtaining the voiceprint features of each sound-making object, the feature identifiers of each part of the audio processing result can be determined based on the voiceprint features of each sound-making object, and the audio processing results can be displayed in partitions on the current interface according to the feature identifiers through a text processing application.

[0138] In some embodiments, it is also possible to determine whether the difference between each voiceprint feature is less than or equal to a set difference threshold, and set the feature identifiers corresponding to the voiceprint features whose differences in the audio processing results are less than or equal to the set difference threshold to be the same.

[0139] In this disclosure, the parts of the audio processing results with the same feature identifiers correspond to the same sounding objects. Taking the audio processing function as an example during a meeting, the audio processing results can be divided into five parts according to the voiceprint features during the implementation process, which are mainly as follows:

[0140] Part 1: The main content of today's meeting is... (Feature Identifier 1)

[0141] Part 2: The problem you raised can be solved using Solution 1. (Feature Identifier 2)

[0142] Part 3: I think Option 2 is also acceptable, because… (Feature ID 3)

[0143] Part 4: Based on the analysis, Solution 2 is more reasonable. (Feature Identifier 1)

[0144] Part 5: I think we can still try Solution 1. (Feature Identifier 2)

[0145] Based on the above content, it can be seen that the feature identifiers of the first part and the fourth part in the audio processing result are both feature identifier 1, then it is determined that the first part and the fourth part correspond to the same sound object (sound object A); the feature identifiers of the second part and the fifth part in the audio processing result are both feature identifier 2, then it is determined that the second part and the fifth part correspond to the same sound object (sound object B); the feature identifier of the third part in the audio processing result is feature identifier 3, then it is determined that the third part corresponds to a sound object (sound object C).

[0146] In the embodiment of the present disclosure, feature identifiers of each part of the audio processing result are determined based on the voiceprint features, and the audio processing result is partitioned and displayed on the current interface according to the feature identifiers through a text processing application. In this way, the various parts of the audio processing result can be presented on the current interface according to the different sound-emitting objects, so that users can classify and identify them.

[0147] In other embodiments, after the various components of the audio processing results are mapped to the sound-producing object based on the feature identifier, a mapping relationship between the voiceprint feature, the sound-producing object, and the feature identifier can be established and stored. In this way, when the mobile terminal restarts the audio processing function, the feature identifiers of each component of the audio processing results can be directly determined based on the above mapping relationship. Compared with processing the audio data again based on the voiceprint processing module, the data processing speed and efficiency can be improved.

[0148] In some embodiments, the mobile terminal includes: a plurality of the audio acquisition modules, and the audio acquisition surfaces of the audio acquisition modules face different directions;

[0149] The acquiring of first audio data by the audio acquisition module based on the mobile terminal includes:

[0150] Based on the multiple audio acquisition modules, multiple groups of second audio data are collected, and the multiple groups of second audio data are fused to obtain the first audio data.

[0151] In an embodiment of the present disclosure, multiple groups of second audio data can be collected based on multiple audio collection modules. After collecting multiple groups of second audio data, the phase difference between each second audio data can be calculated, and the each second audio data can be fused based on the phase difference to obtain the second audio data.

[0152] Since the orientations of the audio collection surfaces of each audio collection module are different, the second audio data can be obtained from various directions and at different angles. Compared with setting up a single audio collection module, the first audio data finally obtained can be more accurate. Moreover, since the first audio data input into the voice interaction application is more accurate, the accuracy of the audio processing results obtained by the voice interaction application can be further improved.

[0153] In some embodiments, the method further comprises:

[0154] In response to an editing instruction for the audio processing result on the text processing application, the audio processing result is edited.

[0155] In an embodiment of the present disclosure, the audio processing results can be output in a text format based on a text processing application. In this way, the user can directly input editing instructions through the text processing application. After the text processing application detects the editing instructions, the audio processing results can be edited in response to the editing instructions.

[0156] For example, in response to a detected modification instruction, the audio processing result in text format can be modified accordingly; in response to a detected addition instruction, input content can be added to the audio processing result in text format; in response to a detected deletion instruction, the audio processing result in text format can be deleted, so that the final generated meeting record is more accurate and complete.

[0157] In some embodiments, at least one type of audio processing module corresponding to the voice interaction application is located in a middle platform device.

[0158] Here, the middle platform device includes: a middle platform access layer with a switching function and a service engine layer at the back end. In the embodiment of the present disclosure, the audio processing function of the mobile terminal can be packaged and integrated through the middle platform device, and the audio processing function can be assigned to the service engine layer through the middle platform access layer. In the embodiment of the present disclosure, by setting at least one type of audio processing module in the service engine layer at the back end of the middle platform device, the processing and calculation functions of a large amount of data are integrated in the service engine layer at the back end of the middle platform device, which can reduce the power consumption of the mobile terminal compared to directly processing and calculating data on the mobile terminal.

[0159] In some embodiments, the method further comprises:

[0160] Sending the authentication information of the mobile terminal to the middle platform device through the communication module, so that the middle platform device authenticates the mobile terminal;

[0161] Wherein, the authentication information at least includes: the device identification of the mobile terminal.

[0162] Here, before the first audio data is processed by the voice interaction application of the mobile terminal, the authentication information is first sent to the middle platform device so that the middle platform device authenticates the mobile terminal to improve the security of the interaction between the mobile terminal and the middle platform device.

[0163] In some embodiments, the processing of the first audio data by the voice interaction application of the mobile terminal to obtain an audio processing result includes:

[0164] When it is determined that the mobile terminal authentication is passed, processing the first audio data by a voice interaction application of the mobile terminal to obtain the audio processing result;

[0165] The mobile terminal authentication includes: a preset identifier in a preset identifier list of the middle station device has a preset identifier that matches the device identifier.

[0166] Here, the device identification may include: the model of the mobile terminal, the International Mobile Equipment Identity (IMEI) code of the mobile terminal, etc., which is used to uniquely identify the mobile terminal.

[0167] During the implementation process, the preset identifiers of mobile terminals with the audio processing function can be pre-stored in the middle station device to form a preset identifier list. During the authentication process, the device identifier of the mobile terminal can be compared with the preset identifiers in the preset representation list to determine whether there is a preset identifier that matches the device identifier.

[0168] Here, the preset identifier that matches the device identifier includes a preset identifier that is identical to the device identifier of the mobile terminal. When the device identifier and the preset identifier are digital codes, if the number of identical digits in the device identifier and the preset identifier exceeds a predetermined threshold, the device identifier and the preset identifier may be determined to match. Of course, other methods may also be used to determine whether the device identifier and the preset identifier match, which are not specifically limited here.

[0169] Since the audio processing function in the present disclosure is set for a specific model of mobile terminal, not all mobile terminals have this audio processing function. During the implementation process, mobile terminals that do not have this audio processing function can be screened out through authentication information to reduce the possibility of invalid interaction and improve the security and confidentiality of using the audio processing function.

[0170] Figure 3 is a flowchart of another audio processing method according to an exemplary embodiment. Figure 3 As shown, the audio processing method is applied to an audio processing system, which includes a mobile terminal and a middleware device. The method mainly includes the following steps:

[0171] In step 301, when the mobile terminal detects an instruction to start the audio processing function, the mobile terminal obtains first audio data based on the audio acquisition module of the mobile terminal;

[0172] In step 302, the mobile terminal obtains an audio processing result obtained by processing the first audio data from the middle platform device through a voice interaction application; wherein the voice interaction application corresponds to at least one type of audio processing module, and the audio processing module is located in the middle platform device, and different types of audio processing modules are used to perform different processing on the first audio data;

[0173] In step 303, the mobile terminal starts a text processing application and outputs the audio processing result through the text processing application.

[0174] Figure 4 FIG. 1 is a schematic diagram showing an architecture of an audio processing system according to an exemplary embodiment. Figure 4 As shown, the audio processing system includes a mobile terminal 401 and a middle platform device 402, and the middle platform device 402 includes: a middle platform access layer 403 with a switching function and a service engine layer 404 at the back end. In some embodiments, a conference simultaneous interpretation access service (first interface) can be integrated in the middle platform access layer, and the first audio data is received through the first interface, and the first audio data is forwarded to the service engine layer through the ASR access service (second interface), and the first audio data is processed by the speech recognition module of the service engine layer. After the speech recognition module obtains the audio processing result, the service engine layer sends the audio processing result to the first interface through the second interface, and forwards the audio processing result to the translation module through the first interface to translate the audio processing result, and sends the obtained translation content to the mobile terminal through the first interface. Among them, the speech recognition module includes: an ASR engine, and the translation module includes: a machine translation engine.

[0175] In some embodiments, the mobile terminal may establish a connection with the middleware device through a preset transmission protocol. For example, the connection with the middleware device may be established based on a full-duplex communication protocol (e.g., the WebSocket protocol) and a secure transport layer protocol (TLS). The WebSocket protocol is a protocol for full-duplex communication over a single Transmission Control Protocol (TCP) connection, and the secure transport layer protocol is used to provide confidentiality and data integrity between two communicating applications.

[0176] In the embodiment of the present disclosure, after the audio processing function of the mobile terminal is turned on, a communication connection between the mobile terminal and the middleware device is established until the audio processing function is turned off.

[0177] Here, the middle platform device includes: a middle platform access layer with a switching function and a service engine layer at the back end. In the embodiment of the present disclosure, the audio processing function of the mobile terminal can be packaged and integrated through the middle platform device, and the audio processing function can be assigned to the service engine layer through the middle platform access layer. In the embodiment of the present disclosure, by setting at least one type of audio processing module in the service engine layer at the back end of the middle platform device, the processing and calculation functions of a large amount of data are integrated in the service engine layer at the back end of the middle platform device, which can reduce the power consumption of the mobile terminal compared to directly processing and calculating data on the mobile terminal.

[0178] In an embodiment of the present disclosure, a mobile terminal can be provided. Since the mobile terminal itself includes an audio acquisition module, a voice interaction application and a text processing application, when the mobile terminal of the present disclosure detects an instruction to start the audio processing function, it can obtain first audio data based on the mobile terminal's own audio acquisition module, and process the first audio data through its own voice interaction application to obtain an audio processing result. After obtaining the audio processing result, it can start its own text processing application and output the audio processing result through the text processing application.

[0179] Since the mobile terminal in the present disclosure can install applications with various functions and has communication functions, it can be used by users in daily life. Users generally carry it with them. In this way, even if a meeting is notified at short notice, as long as the user carries the mobile terminal, he or she only needs to turn on the audio processing function of the mobile terminal during the meeting, and can call the mobile terminal's own audio acquisition module to perform audio acquisition, and process the collected first audio data through the voice interaction application to obtain an audio processing result, and finally output the audio processing result through a text processing application, without the need to temporarily prepare additional recording and video equipment, which provides convenience for the formation of meeting minutes.

[0180] In some embodiments, the mobile terminal obtains, from the middle station device, an audio processing result obtained by processing the first audio data through a voice interaction application, including:

[0181] The mobile terminal sends authentication information to the middle station device through the communication module; wherein the authentication information includes: the device identification of the mobile terminal;

[0182] The middle station device compares the device identifier with the preset identifiers in the preset identifier list to determine whether there is a preset identifier in the preset identifier list that matches the device identifier;

[0183] When the middle station device determines that there is a preset identifier matching the device identifier in the preset identifier list, it determines that the mobile terminal authentication is successful.

[0184] Here, the device identification may include: the model of the mobile terminal, the IMEI code of the mobile terminal, etc., which is used to uniquely identify the mobile terminal.

[0185] During the implementation process, the preset identifiers of mobile terminals with the audio processing function can be pre-stored in the middle station device to form a preset identifier list. During the authentication process, the device identifier of the mobile terminal can be compared with the preset identifiers in the preset representation list to determine whether there is a preset identifier that matches the device identifier.

[0186] Here, the preset identifier that matches the device identifier includes a preset identifier that is identical to the device identifier of the mobile terminal. When the device identifier and the preset identifier are digital codes, if the number of identical digits in the device identifier and the preset identifier exceeds a predetermined threshold, the device identifier and the preset identifier may be determined to match. Of course, other methods may also be used to determine whether the device identifier and the preset identifier match, which are not specifically limited here.

[0187] Since the audio processing function in the present disclosure is set for a specific model of mobile terminal, not all mobile terminals have this audio processing function. During the implementation process, mobile terminals that do not have this audio processing function can be screened out through authentication information to reduce the possibility of invalid interaction.

[0188] Figure 5 FIG. 1 is a structural block diagram of a mobile terminal according to an exemplary embodiment. Figure 5 As shown, the mobile terminal 500 mainly includes:

[0189] The audio acquisition module 501 is configured to acquire first audio data when detecting an activation instruction for the audio processing function of the mobile terminal;

[0190] A voice interaction application 502 is configured to process the first audio data to obtain an audio processing result; wherein the voice interaction application corresponds to at least one type of audio processing module, and the different types of audio processing modules are respectively used to perform different processing on the first audio data;

[0191] The text processing application 503 is configured to output the audio processing result.

[0192] In some embodiments, the text processing application 503 is further configured to:

[0193] The audio processing result is output in text format.

[0194] In some embodiments, the at least one type of audio processing module includes: a speech recognition module; the speech recognition module is configured to:

[0195] The first audio data is recognized by the voice recognition module of the voice interaction application to obtain semantic content corresponding to the first audio data.

[0196] In some embodiments, the at least one type of audio processing module further includes: a translation module; the translation module is configured to:

[0197] When it is determined that the translation function of the mobile terminal is turned on, the semantic content is translated by the translation module of the voice interaction application to obtain translation content expressed in a target language.

[0198] In some embodiments, the at least one type of audio processing module further includes: a voiceprint processing module; the voiceprint processing module is configured to:

[0199] Obtain the voiceprint characteristics of each sound-making object;

[0200] Determining feature identifiers of various parts of the audio processing result based on the voiceprint features of each of the sound-making objects;

[0201] The text processing application is further configured to display the audio processing results in partitions on the current interface according to the feature identifiers.

[0202] In some embodiments, the mobile terminal 500 includes: a plurality of the audio acquisition modules, and the audio acquisition surfaces of the audio acquisition modules face different directions;

[0203] The plurality of audio acquisition modules are further configured as follows:

[0204] Collect multiple sets of second audio data, and perform fusion processing on the multiple sets of second audio data to obtain the first audio data.

[0205] In some embodiments, the text processing application 503 is further configured to:

[0206] In response to an editing instruction for the audio processing result, the audio processing result is edited.

[0207] In some embodiments, the mobile terminal 500 further includes:

[0208] a communication module configured to send the authentication information of the mobile terminal to the middle platform device so that the middle platform device authenticates the mobile terminal;

[0209] Wherein, the authentication information at least includes: the device identification of the mobile terminal.

[0210] In some embodiments, the voice interaction application 502 is further configured to:

[0211] When it is determined that the mobile terminal authentication is passed, processing the first audio data to obtain the audio processing result;

[0212] The mobile terminal authentication includes: a preset identifier in a preset identifier list of the middle station device has a preset identifier that matches the device identifier.

[0213] In some embodiments, at least one type of audio processing module corresponding to the voice interaction application 502 is located in the middle platform device.

[0214] Figure 6 FIG. 1 is a block diagram of an audio processing system according to an exemplary embodiment. Figure 6 As shown, the audio processing system 600 mainly includes:

[0215] Mobile terminal 601 and middleware 602;

[0216] The mobile terminal 601 is configured to acquire first audio data based on the audio acquisition module of the mobile terminal when detecting an activation instruction for the audio processing function;

[0217] Obtaining, from the middle platform device 602, an audio processing result obtained by processing the first audio data through a voice interaction application; wherein the voice interaction application corresponds to at least one type of audio processing module, the audio processing module being located in the middle platform device, and the different types of audio processing modules being used to perform different processing on the first audio data;

[0218] The mobile terminal 601 is further configured to start a text processing application and output the audio processing result through the text processing application.

[0219] In some embodiments, the mobile terminal 601 is further configured to send authentication information to the middle station device 602 through a communication module; wherein the authentication information includes: a device identifier of the mobile terminal;

[0220] The middle station device 602 is configured to compare the device identifier with a preset identifier in a preset identifier list to determine whether there is a preset identifier in the preset identifier list that matches the device identifier;

[0221] When it is determined that there is a preset identifier matching the device identifier in the preset identifier list, it is determined that the authentication of the mobile terminal 601 is successful.

[0222] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0223] Figure 7 1 is a block diagram showing a mobile terminal 1200 according to an exemplary embodiment. For example, the mobile terminal 1200 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0224] Reference Figure 7 The mobile terminal 1200 may include one or more of the following components: a processing component 1202 , a memory 1204 , a power component 1206 , a multimedia component 1208 , an audio component 1210 , an input / output (I / O) interface 1212 , a sensor component 1214 , and a communication component 1216 .

[0225] The processing component 1202 generally controls the overall operation of the mobile terminal 1200, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 1202 may include one or more processors 1220 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 1202 may include one or more modules to facilitate interaction between the processing component 1202 and other components. For example, the processing component 1202 may include a multimedia module to facilitate interaction between the multimedia component 1208 and the processing component 1202.

[0226] The memory 1204 is configured to store various types of data to support the operations of the device 1200. Examples of such data include instructions for any application or method operating on the mobile terminal 1200, contact data, phone book data, messages, pictures, videos, etc. The memory 1204 can be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0227] The power component 1206 provides power to various components of the mobile terminal 1200. The power component 1206 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the mobile terminal 1200.

[0228] The multimedia component 1208 includes a screen that provides an output interface between the mobile terminal 1200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 1208 includes a front camera and / or a rear camera. When the device 1200 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.

[0229] The audio component 1210 is configured to output and / or input audio signals. For example, the audio component 1210 includes a microphone (MIC), which is configured to receive external audio signals when the mobile terminal 1200 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 1204 or transmitted via the communication component 1216. In some embodiments, the audio component 1210 also includes a speaker for outputting audio signals.

[0230] I / O interface 1212 provides an interface between processing component 1202 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include but are not limited to: a home button, volume buttons, a start button, and a lock button.

[0231] Sensor assembly 1214 includes one or more sensors for providing various aspects of the status assessment of mobile terminal 1200. For example, sensor assembly 1214 can detect the open / closed state of device 1200, the relative positioning of components, such as the display and keypad of mobile terminal 1200. Sensor assembly 1214 can also detect changes in the position of mobile terminal 1200 or a component of mobile terminal 1200, the presence or absence of user contact with mobile terminal 1200, the orientation or acceleration / deceleration of mobile terminal 1200, and changes in the temperature of mobile terminal 1200. Sensor assembly 1214 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 1214 can also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 1214 can also include an accelerometer, a gyroscope, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0232] The communication component 1216 is configured to facilitate wired or wireless communication between the mobile terminal 1200 and other devices. The mobile terminal 1200 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 1216 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 1216 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0233] In an exemplary embodiment, the mobile terminal 1200 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.

[0234] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 1204 including instructions, which can be executed by the processor 1220 of the mobile terminal 1200 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0235] A non-transitory computer-readable storage medium, when instructions in the storage medium are executed by a processor of a mobile terminal, enables the mobile terminal to perform an audio processing method, the method comprising:

[0236] When an activation instruction for the audio processing function of the mobile terminal is detected, acquiring first audio data based on an audio acquisition module of the mobile terminal;

[0237] Processing the first audio data by the voice interaction application of the mobile terminal to obtain an audio processing result; wherein the voice interaction application corresponds to at least one type of audio processing module, and the different types of audio processing modules are respectively used to perform different processing on the first audio data;

[0238] The text processing application of the mobile terminal is started, and the audio processing result is output through the text processing application.

[0239] Figure 81 is a block diagram of a middleware device 1300 according to an exemplary embodiment. For example, the device 1300 may be provided as a server. Figure 8 , device 1300 includes a processing component 1322, which further includes one or more processors, and memory resources represented by memory 1332 for storing instructions, such as applications, that are executable by the processing component 1322. The applications stored in memory 1332 may include one or more modules, each corresponding to a set of instructions.

[0240] The device 1300 may also include a power supply component 1326 configured to perform power management of the device 1300, a wired or wireless network interface 1350 configured to connect the device 1300 to a network, and an input / output (I / O) interface 1358. The device 1300 may operate based on an operating system stored in the memory 1332, such as Windows Server™, MacOS X™, Unix™, Linux™, FreeBSD™, or the like.

[0241] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.

[0242] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. An audio processing method, characterized in that: The method comprises: When the mobile terminal detects an instruction to activate an audio processing function of the mobile terminal, the mobile terminal enters a conference recording mode. In the conference recording mode, multiple sets of second audio data are collected based on multiple audio collection modules of the mobile terminal, the multiple sets of second audio data are fused and processed to obtain first audio data, and the first audio data is transmitted to a voice interaction application installed in the mobile terminal through a first connection path; wherein the audio collection surfaces of the respective audio collection modules face different directions; When the voice interaction application receives the first audio data, it processes the first audio data through the audio processing module to obtain an audio processing result, and transmits the audio processing result to the text processing application installed in the mobile terminal through the second connection path; wherein, the voice interaction application corresponds to at least one type of audio processing module, and the audio processing module is located in the middle platform device, and different types of audio processing modules are used to perform different processing on the first audio data; the middle platform device includes: a middle platform access layer with a switching function and a service engine layer located at the back end, the middle platform access layer receives the first audio data through the first interface, forwards the first audio data to the service engine layer through the second interface, and processes the first audio data through the voice recognition module of the service engine layer to obtain the audio processing result; after the voice recognition module obtains the audio processing result, the service engine layer sends the audio processing result to the first interface through the second interface, and forwards the audio processing result to the translation module through the first interface to translate the audio processing result to obtain the translation content expressed in the target language, and sends the obtained translation content to the mobile terminal through the first interface; The text processing application of the mobile terminal is started, and the translated content is outputted while the audio processing result is outputted in a text format through the text processing application.

2. The method according to claim 1, characterized in that The processing of the first audio data by the audio processing module to obtain an audio processing result includes: The first audio data is recognized by the voice recognition module of the voice interaction application to obtain semantic content corresponding to the first audio data.

3. The method according to claim 2, characterized in that The step of processing the first audio data by the audio processing module to obtain an audio processing result further includes: When it is determined that the translation function of the mobile terminal is turned on, the semantic content is translated by the translation module of the voice interaction application to obtain translation content expressed in a target language.

4. The method according to claim 1, wherein The at least one type of audio processing module further includes: a voiceprint processing module; the method further includes: Inputting the first audio data into the voiceprint processing module of the voice interaction application to obtain the voiceprint features of each sound-making object; Determining feature identifiers of various parts of the audio processing result based on the voiceprint features of each of the sound-making objects; The text processing application displays the audio processing results in partitions on the current interface according to the feature identifiers.

5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: In response to an editing instruction for the audio processing result on the text processing application, the audio processing result is edited.

6. The method according to any one of claims 1 to 4, characterized in that The method further comprises: Sending the authentication information of the mobile terminal to the middle platform device through the communication module, so that the middle platform device authenticates the mobile terminal; Wherein, the authentication information at least includes: the device identification of the mobile terminal.

7. The method according to claim 6, characterized in that The processing of the first audio data by the audio processing module to obtain an audio processing result includes: When it is determined that the mobile terminal authentication is passed, processing the first audio data by a voice interaction application of the mobile terminal to obtain the audio processing result; The mobile terminal authentication includes: a preset identifier in a preset identifier list of the middle station device has a preset identifier that matches the device identifier.

8. An audio processing method, characterized in that: The method comprises: The audio processing system includes a mobile terminal that enters a conference recording mode when detecting an instruction to activate an audio processing function. In the conference recording mode, multiple audio collection modules of the mobile terminal collect multiple sets of second audio data, perform fusion processing on the multiple sets of second audio data to obtain first audio data, and transmit the first audio data to a voice interaction application installed in the mobile terminal through a first connection path; wherein the audio collection surfaces of the audio collection modules are oriented in different directions; The mobile terminal obtains the audio processing result obtained by processing the first audio data from the middle platform device included in the audio processing system through the voice interaction application, and transmits the audio processing result to the text processing application installed in the mobile terminal through the second connection path; wherein, the voice interaction application corresponds to at least one type of audio processing module, the audio processing module is located in the middle platform device, and different types of audio processing modules are used to perform different processing on the first audio data; the middle platform device includes: a middle platform access layer with a switching function and a service engine layer located at the back end, the middle platform access layer receives the first audio data through a first interface, forwards the first audio data to the service engine layer through a second interface, and processes the first audio data through the voice recognition module of the service engine layer to obtain the audio processing result; after the voice recognition module obtains the audio processing result, the service engine layer sends the audio processing result to the first interface through the second interface, and forwards the audio processing result to the translation module through the first interface to translate the audio processing result to obtain the translation content expressed in the target language, and sends the obtained translation content to the mobile terminal through the first interface; The mobile terminal starts a text processing application and outputs the translated content while outputting the audio processing result in a text format through the text processing application.

9. The method according to claim 8, characterized in that The mobile terminal obtains, through the voice interaction application, an audio processing result obtained by processing the first audio data from the middle station device, including: The mobile terminal sends authentication information to the middle station device through the communication module; wherein the authentication information includes: the device identification of the mobile terminal; The middle station device compares the device identifier with the preset identifiers in the preset identifier list to determine whether there is a preset identifier in the preset identifier list that matches the device identifier; When the middle station device determines that there is a preset identifier matching the device identifier in the preset identifier list, it determines that the mobile terminal authentication is successful.

10. A mobile terminal, characterized in that: include: Multiple audio collection modules are configured to enter a conference recording mode upon detecting an activation instruction for an audio processing function of the mobile terminal, collect multiple sets of second audio data in the conference recording mode, and fuse the multiple sets of second audio data to obtain first audio data, and the mobile terminal transmits the first audio data to a voice interaction application installed in the mobile terminal via a first connection path; wherein the audio collection surfaces of the audio collection modules are oriented in different directions; A voice interaction application is configured to, upon receiving the first audio data, process the first audio data through an audio processing module to obtain an audio processing result, and transmit the audio processing result to a text processing application installed in the mobile terminal through a second connection path; wherein the voice interaction application corresponds to at least one type of audio processing module, the audio processing module is located in a middle platform device, and different types of audio processing modules are used to perform different processing on the first audio data; the middle platform device includes: a middle platform access layer with a switching function and a service engine layer located at the back end, the middle platform access layer receives the first audio data through a first interface, forwards the first audio data to the service engine layer through a second interface, and processes the first audio data through a voice recognition module of the service engine layer to obtain the audio processing result; after the voice recognition module obtains the audio processing result, the service engine layer sends the audio processing result to the first interface through the second interface, and forwards the audio processing result to the translation module through the first interface to translate the audio processing result to obtain translation content expressed in the target language, and sends the obtained translation content to the mobile terminal through the first interface; The text processing application is configured to output the translated content while outputting the audio processing result in a text format through the text processing application.

11. The mobile terminal according to claim 10, wherein: The speech recognition module is configured as follows: The first audio data is recognized by the voice recognition module of the voice interaction application to obtain semantic content corresponding to the first audio data.

12. The mobile terminal according to claim 11, wherein: The translation module is configured as follows: When it is determined that the translation function of the mobile terminal is turned on, the semantic content is translated by the translation module of the voice interaction application to obtain translation content expressed in a target language.

13. The mobile terminal according to claim 10, wherein: The at least one type of audio processing module further includes: a voiceprint processing module; the voiceprint processing module is configured to: Obtain the voiceprint characteristics of each sound-making object; Determining feature identifiers of various parts of the audio processing result based on the voiceprint features of each of the sound-making objects; The text processing application is further configured to display the audio processing results in partitions on the current interface according to the feature identifiers.

14. The mobile terminal according to any one of claims 10 to 13, characterized in that: The text processing application is further configured to: In response to an editing instruction for the audio processing result, the audio processing result is edited.

15. The mobile terminal according to any one of claims 10 to 13, characterized in that: The mobile terminal further includes: a communication module configured to send the authentication information of the mobile terminal to the middle platform device so that the middle platform device authenticates the mobile terminal; Wherein, the authentication information at least includes: the device identification of the mobile terminal.

16. The mobile terminal according to claim 15, characterized in that: The voice interaction application is further configured as follows: When it is determined that the mobile terminal authentication is passed, processing the first audio data to obtain the audio processing result; The mobile terminal authentication includes: a preset identifier in a preset identifier list of the middle station device has a preset identifier that matches the device identifier.

17. An audio processing system, characterized in that: include: Mobile terminals and middleware devices; The mobile terminal is configured to enter a conference recording mode upon detecting an activation instruction for an audio processing function, collect multiple sets of second audio data based on multiple audio collection modules of the mobile terminal in the conference recording mode, perform fusion processing on the multiple sets of second audio data to obtain first audio data, and transmit the first audio data to a voice interaction application installed in the mobile terminal via a first connection path; wherein the audio collection surfaces of the respective audio collection modules face different directions; Through the voice interaction application, an audio processing result obtained by processing the first audio data is obtained from the middle platform device, and the audio processing result is transmitted to the text processing application installed in the mobile terminal through the second connection path; wherein, the voice interaction application corresponds to at least one type of audio processing module, the audio processing module is located in the middle platform device, and different types of audio processing modules are used to perform different processing on the first audio data; the middle platform device includes: a middle platform access layer with a switching function and a service engine layer located at the back end, the middle platform access layer receives the first audio data through a first interface, forwards the first audio data to the service engine layer through a second interface, and processes the first audio data through the voice recognition module of the service engine layer to obtain the audio processing result; after the voice recognition module obtains the audio processing result, the service engine layer sends the audio processing result to the first interface through the second interface, and forwards the audio processing result to the translation module through the first interface to translate the audio processing result to obtain the translation content expressed in the target language, and sends the obtained translation content to the mobile terminal through the first interface; The mobile terminal is further configured to start a text processing application and output the translated content while outputting the audio processing result in a text format through the text processing application.

18. The audio processing system according to claim 17, wherein: The mobile terminal is further configured to send authentication information to the middle station device through the communication module; wherein the authentication information includes: a device identifier of the mobile terminal; The middle station device is configured to compare the device identifier with a preset identifier in a preset identifier list to determine whether there is a preset identifier in the preset identifier list that matches the device identifier; When it is determined that there is a preset identifier in the preset identifier list that matches the device identifier, it is determined that the mobile terminal authentication is successful.

19. A mobile terminal, characterized in that: include: processor; a memory configured to store processor-executable instructions; The processor is configured to implement the steps of any one of the audio processing methods in claims 1 to 9 when executing.

20. A non-transitory computer-readable storage medium, when instructions in the storage medium are executed by a processor of a mobile terminal, enables the mobile terminal to perform any one of the audio processing methods of claims 1 to 9.

Citation Information

Patent Citations

  • Meeting record generation method, apparatus, apparatus, and computer storage medium

    CN109388701A

  • Call information processing method and device and computer storage medium

    CN111510556A

  • Voice interaction method, voice interaction equipment, computing equipment and storage medium

    CN112449050A