Method and apparatus for processing audio
By identifying and adjusting the speaking style characteristics of the audio stream, and utilizing coaching and recognition models, adjustment prompts and processing are provided, which solves the problem of insufficient audio guidance for electronic devices in live streaming, and improves audio processing capabilities and user experience.
Patent Information
- Application Number
- CN202111172819.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-08
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2041-10-08
AI Technical Summary
Existing electronic devices lack audio guidance for live streamers, resulting in poor live streaming quality and unsatisfactory user experience.
By acquiring the audio stream, identifying speech pattern characteristics, and calling the coaching model and speech pattern recognition model, we can provide adjustment prompts and processing for voice, tone, and speech rate, and optimize the audio stream to improve audio functionality.
It improves the audio processing capabilities of electronic devices during live streaming, enhances the accuracy of user speech recognition and audio quality, and meets users' personalized needs.
Smart Images

Figure CN115966198B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of electronic information, and in particular, to an audio processing method and device. BACKGROUND
[0002] With the development of information technology, live broadcast, an online real-time interaction mode using multimedia as a medium, has been widely applied in various industries, such as sales, training, and performance industries.
[0003] Therefore, users have higher and higher requirements for the audio function of electronic devices. However, the audio function of the electronic device still has room for improvement. SUMMARY
[0004] The present application provides an audio processing method and device, aiming to solve the problem of how to improve the audio function of electronic devices.
[0005] To achieve the above-mentioned purpose, the present application provides the following technical solutions:
[0006] The first aspect of the present application provides an audio processing method applied to an electronic device, comprising: obtaining an audio stream, in response to a speaking manner recognition result of the audio stream indicating that the speaking manner needs to be adjusted, displaying a speaking manner adjustment prompt information, and / or outputting an audio stream after speaking manner adjustment. Displaying the speaking manner adjustment prompt information and outputting the audio stream after speaking manner adjustment can both achieve the adjustment of the speaking manner in the case that the speaking manner recognition result indicates that the speaking manner needs to be adjusted, thereby improving the audio function of the electronic device.
[0007] Optionally, before the response to the speaking manner recognition result of the audio stream indicating that the speaking manner needs to be adjusted, displaying the speaking manner adjustment prompt information, and / or outputting the audio stream after speaking manner adjustment, the method further comprises: extracting speaking manner features of the audio stream, the speaking manner features comprising at least one of voice features, tone features, and speed features of the audio stream, and obtaining the speaking manner recognition result of the audio stream according to the speaking manner features of the audio stream. According to at least one of the voice features, the tone features, and the speed features of the audio stream, the accuracy of the obtained speaking manner recognition result of the audio stream can be improved.
[0008] Optionally, the obtaining of the speech manner recognition result of the audio stream according to the speech manner feature of the audio stream comprises: calling at least one of a target coach model and a speech manner recognition model to obtain the speech manner recognition result of the audio stream according to the speech manner feature of the audio stream, the target coach model being a model configured according to at least one of an audio effect expected by an audio input party and attribute information of the audio input party. The speech manner recognition result of the audio stream is determined from at least one of two dimensions, and the target coach model is configured according to at least one of the two dimensions, which can further improve the accuracy.
[0009] Optionally, the configuration process of the target coach model comprises: in response to a selection instruction, displaying a coach model selection interface, the coach model selection interface being used to receive at least one of an audio effect expected by an audio input party and attribute information of the audio input party, in response to an operation in the coach model selection interface, displaying a coach model recommendation interface, and in response to an operation in the coach model recommendation interface, confirming the selected target coach model. The way of obtaining the target coach model based on the interactive interface has higher user experience and implementability.
[0010] Optionally, the display of the coach model selection interface comprises: displaying a first interface and / or a second interface, the first interface comprising: an industry option, a gender option and a character model option, and the second interface comprising: a receiving control of at least one of an audio sample and a text sample. The division of the first interface and the second interface is conducive to showing clear selection logic to the user to obtain better user experience.
[0011] Optionally, the coach model selection interface further comprises a third interface; before the display of the coach model recommendation interface in response to the operation in the coach model selection interface, the method further comprises: in response to a conflict between the audio effect expected by the audio input party and the attribute information of the audio input party, displaying the third interface, the third interface being used to receive secondary confirmation information of the conflicting information to further improve the fitting degree of the target coach model and the effect expected by the user.
[0012] Optionally, the display of the coach model recommendation interface comprises: displaying the coach model recommendation interface comprising information of recommended coach models, the information of a first recommended coach model of any one of the recommended coach models comprising: an icon control and summary information, the icon control being used to trigger the playing of an audio of the first recommended coach model, and the summary information comprising: an introduction of at least one of a category and a style to which a voice, a speech rate and a tone of the first recommended coach model belong, so as to show more comprehensive information of the coach model and facilitate the user to perceive the information of the coach model.
[0013] Optionally, the speaking manner adjustment prompt information comprises at least one of adjustment of voice, tone and speech speed; and the obtaining manner of the audio stream after speaking manner adjustment comprises processing at least one of voice, tone and speech speed of the audio stream to obtain the audio stream after speaking manner adjustment. At least one of voice, tone and speech speed of the audio stream can reflect speaking manner, so that adjusting at least one of voice, tone and speech speed of the audio stream can better achieve adjustment of speaking manner.
[0014] Optionally, before the displaying of the speaking manner adjustment prompt information and / or the outputting of the audio stream after speaking manner adjustment, the method further comprises: dividing the first type of features and the second type of features in the speaking manner features; the second type of features is more difficult to be manually adjusted than the first type of features; the displaying of the speaking manner adjustment prompt information comprises displaying adjustment prompt information of the first type of features; and the outputting of the audio stream after speaking manner adjustment comprises outputting the audio stream after adjustment of the second type of features. Displaying adjustment prompt information of features easy to be manually adjusted and processing features difficult to be manually adjusted can be beneficial to balancing processing resources and adjustment effect.
[0015] Optionally, the method further comprises: in response to an audio collection end instruction, displaying a feedback interface, the feedback interface being configured to receive feedback information, the feedback information comprising evaluation information of each audio frame in the output audio stream, so as to collect feedback information of adjustment effect.
[0016] Optionally, the method further comprises: adjusting the speaking manner recognition model according to the speaking manner features of the audio frame whose evaluation information meets a condition, so as to optimize the speaking manner recognition model and improve accuracy of the speaking manner recognition result.
[0017] A second aspect of the present application provides an electronic device, comprising a display screen, a processor and a memory; the memory is configured to store an application program, and the processor is configured to run the application program to implement the audio processing method provided by the first aspect of the present application.
[0018] A third aspect of the present application provides a readable storage medium having an application program stored thereon, and the audio processing method provided by the first aspect of the present application is implemented when a computer device runs the application program. BRIEF DESCRIPTION OF DRAWINGS
[0019] Figure 1 An example diagram of an online live scene;
[0020] Figure 2 An example diagram of a structure of an electronic device disclosed by an embodiment of the present application;
[0021] Figure 3An example diagram of a software framework of an electronic device disclosed in an embodiment of the present application;
[0022] Figure 4 An example diagram of a flowchart of a method for processing audio disclosed in an embodiment of the present application;
[0023] Figure 5a An example diagram of a first interface of a coach model selection interface disclosed in an embodiment of the present application;
[0024] Figure 5b An example diagram of a second interface of a coach model selection interface disclosed in an embodiment of the present application;
[0025] Figure 5c An example diagram of a third interface of a coach model selection interface disclosed in an embodiment of the present application;
[0026] Figure 5d An example diagram of a coach model recommendation interface disclosed in an embodiment of the present application;
[0027] Figure 6 An example diagram of displaying speaking manner adjustment prompt information disclosed in an embodiment of the present application;
[0028] Figure 7 An example diagram of a feedback interface disclosed in an embodiment of the present application;
[0029] Figure 8 An example diagram of a flowchart of another method for processing audio disclosed in an embodiment of the present application. DETAILED DESCRIPTION
[0030] Figure 1 An example of a live online scenario:
[0031] A user collects a video through an electronic device, and transmits the video to other devices through a network.
[0032] Audio is an important component of a video. Parameters of audio include: voice (such as technical parameters including but not limited to pitch, volume, length, and tone), speed (such as technical parameters including but not limited to length of syllable and tightness of connection), and intonation (such as technical parameters including but not limited to tone, tone, and pause).
[0033] Different combinations of voice, speed, and intonation will bring different feelings to the listener: for example, soft tone and moderate syllable length will bring comfortable feelings to the listener, while too high pitch and too few pauses will cause discomfort to the listener, and for another example, too flat tone will cause the audio to have no characteristics and be insufficient to attract the listener.
[0034] Most users who perform live streaming have not undergone professional training in pronunciation, so it is possible to cause the listener to have poor feelings.
[0035] However, the existing electronic device lacks guidance or processing of the audio of the live streamer for the purpose of improving the live streaming effect, and therefore the audio processing function of the electronic device needs to be improved.
[0036] To solve the above problems, the embodiments of the present application disclose an audio processing method applied to an electronic device.
[0037] In some embodiments, the electronic device can be a mobile phone, a tablet computer, a desktop computer, a laptop computer, a notebook computer, an ultra-mobile personal computer (UMPC), a handheld computer, a netbook, a personal digital assistant (PDA), a wearable electronic device, a smart watch, etc. The specific form of the smart home device, the server and the electronic device is not specially limited in the present application. In the present embodiment, the electronic device is taken as an example of a mobile phone, and the structure can be as shown in Figure 2 The mobile phone can include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charge management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a loudspeaker 170A, a receiver 170B, a microphone 170C, a headset interface 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 can include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0038] It can be understood that the structure shown in the present embodiment does not constitute a specific limitation on the electronic device. In other embodiments, the electronic device can include more or fewer components than those shown, or combine certain components, or split certain components, or different arrangement of components. The components shown can be implemented in hardware, software or a combination of software and hardware.
[0039] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units can be independent devices, or can be integrated in one or more processors. For example, in the present application, the processor 110 can extract the features of the audio, and use the model to identify the features of the audio to obtain the identification result, and process or provide adjustment prompt information according to the identification result.
[0040] Among them, the controller can be the nerve center and command center of the electronic device. The controller can generate operation control signals according to instruction operation codes and timing signals to complete the control of instruction fetching and instruction execution.
[0041] The processor 110 can also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory can save instructions or data that the processor 110 has just used or repeatedly uses. If the processor 110 needs to use the instructions or data again, it can directly call from the memory. Avoiding repeated access, reducing the waiting time of the processor 110, thus improving the efficiency of the system.
[0042] In some embodiments, the processor 110 can include one or more interfaces. The interfaces can include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0043] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 can include multiple sets of I2C buses. The processor 110 can be coupled to the touch sensor 180K, the charger, the flash, the camera 193, etc. through different I2C bus interfaces, respectively. For example, the processor 110 can be coupled to the touch sensor 180K through an I2C interface, so that the processor 110 and the touch sensor 180K communicate through the I2C bus interface, and realize the touch function of the electronic device.
[0044] The I2S interface can be used for audio communication. In some embodiments, the processor 110 can include multiple sets of I2S buses. The processor 110 can be coupled to the audio module 170 through the I2S bus, and realize the communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 through the I2S interface, and realize the function of answering the phone through the Bluetooth earphone.
[0045] The PCM interface can also be used for audio communication, which samples, quantizes and encodes analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled through the PCM bus interface. In some embodiments, the audio module 170 can also transmit audio signals to the wireless communication module 160 through the PCM interface, and realize the function of answering the phone through the Bluetooth earphone. Both the I2S interface and the PCM interface can be used for audio communication.
[0046] The UART interface is a universal serial bus for asynchronous communication. The bus can be a bidirectional communication bus. It converts the data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is usually used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 through the UART interface to realize the Bluetooth function. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 through the UART interface to realize the function of playing music through the Bluetooth headset.
[0047] The MIPI interface can be used to connect the processor 110 and peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a camera serial interface (CSI), a display serial interface (DSI), etc. In some embodiments, the processor 110 and the camera 193 communicate through the CSI interface to realize the shooting function of the electronic device. The processor 110 and the display screen 194 communicate through the DSI interface to realize the display function of the electronic device.
[0048] The GPIO interface can be configured by software. The GPIO interface can be configured as a control signal or as a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 and the camera 193, the display screen 194, the wireless communication module 160, the audio module 170, the sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.
[0049] The USB interface 130 is an interface that meets the USB standard specification, which can be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 130 can be used to connect a charger to charge the electronic device, or to transmit data between the electronic device and a peripheral device. It can also be used to connect a headset to play audio through the headset. The interface can also be used to connect other electronic devices, such as AR devices, etc.
[0050] It can be understood that the interface connection relationship between the modules illustrated in the embodiments is only illustrative and does not constitute a structural limitation of the electronic device. In some other embodiments of the present application, the electronic device can also use different interface connection methods or combinations of multiple interface connection methods in the above embodiments.
[0051] The electronic device implements a display function through a GPU, the display 194, and an application processor, etc. The GPU is a microprocessor for image processing, connected to the display 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 can include one or more GPUs that execute program instructions to generate or change display information.
[0052] The display 194 is used to display images, videos, etc. The display 194 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flex light-emitting diode (FLED), a Miniled, a Micro Led, a Micro-oled, a quantum dot light emitting diode (QLED), etc. In some embodiments, the electronic device can include 1 or N display screens 194, N being a positive integer greater than 1.
[0053] A series of graphical user interfaces (GUIs) can be displayed on the display 194 of the electronic device, and these GUIs are all home screens of the electronic device. Generally, the size of the display 194 of the electronic device is fixed, and only limited controls can be displayed in the display 194 of the electronic device. A control is a GUI element, which is a software component included in an application program, controls all data processed by the application program and interactive operations related to the data, and a user can interact with the control through direct manipulation to read or edit relevant information of the application program. Generally, a control can include an icon, a button, a menu, a tab, a text box, a dialog box, a status bar, a navigation bar, a widget, and other visual interface elements.
[0054] The electronic device can implement a shooting function through an ISP, the camera 193, a video codec, a GPU, the display 194, and an application processor, etc.
[0055] ISP is used to process the data feedback from the camera 193. For example, when taking a photo, the shutter is opened, the light is transmitted to the camera photosensitive element through the lens, the light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and conversion into a visible image. ISP can also optimize the noise, brightness, and skin color of the image. ISP can also optimize the exposure, color temperature, and other parameters of the shooting scene. In some embodiments, ISP can be provided in the camera 193.
[0056] The camera 193 is used to capture still images or videos. Objects generate optical images through lenses and project them onto photosensitive elements. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into a standard RGB, YUV, or other format image signal. In some embodiments, the electronic device can include one or N cameras 193, where N is a positive integer greater than 1.
[0057] The digital signal processor is used to process digital signals, in addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.
[0058] The video codec is used to compress or decompress digital video. The electronic device can support one or more video codecs. In this way, the electronic device can play or record videos in multiple encoding formats, such as moving picture experts group (MPEG) 1, MPEG 2, MPEG 3, MPEG 4, etc.
[0059] NPU is a neural-network (NN) computing processor that learns from biological neural network structures, such as the transmission mode between human brain neurons, to quickly process input information and continuously self-learn. Through NPU, the electronic device can achieve intelligent cognition applications, such as image recognition, face recognition, voice recognition, text understanding, etc.
[0060] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.
[0061] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of the electronic device by running the instructions stored in internal memory 121. For example, in this embodiment, processor 110 can perform scene arrangement by executing the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc. The data storage area may store data created during the use of the electronic device (such as audio data, phone book, etc.). In addition, internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 110 executes various functional applications and data processing of the electronic device by running the instructions stored in internal memory 121 and / or instructions stored in memory disposed in the processor.
[0062] Electronic devices can implement audio functions such as music playback and recording through audio modules 170, speakers 170A, receivers 170B, microphones 170C, headphone jacks 170D, and application processors.
[0063] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.
[0064] The speaker 170A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. Electronic devices can listen to music or make hands-free calls through the speaker 170A.
[0065] The receiver 170B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When an electronic device answers a phone call or voice message, the receiver 170B can be brought close to the ear to hear the voice.
[0066] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 170C, inputting the sound signal into microphone 170C. Electronic devices can have at least one microphone 170C. In some embodiments, electronic devices can have two microphones 170C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, electronic devices can have three, four, or more microphones 170C, enabling sound signal collection, noise reduction, sound source identification, and directional recording, among other functions.
[0067] An operating system runs on top of the aforementioned components. For example... operating system, Open source operating systems Operating system, etc. Applications can be installed and run on this operating system.
[0068] The operating system of an electronic device can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application's embodiment uses a layered architecture. Taking the system as an example, the software structure of the electronic device is illustrated.
[0069] Figure 3 This is a software structure block diagram of an electronic device according to an embodiment of this application.
[0070] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, [the following is omitted as the text is incomplete and likely refers to a specific implementation or feature]. The system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.
[0071] The application layer can include a series of application packages. For example... Figure 3 As shown, the application package may include applications such as camera, gallery, and video. In this embodiment, the application package may also include live streaming for providing live streaming functionality.
[0072] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions. For example... Figure 3As shown, the application framework layer may include a window manager, content provider, view system, resource manager, notification manager, etc. For example, in this embodiment, the application framework layer can provide APIs related to live streaming functionality to the application layer, and provide live streaming interface management services to the application layer to implement live streaming functionality.
[0073] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.
[0074] Content providers store and retrieve data, making that data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.
[0075] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.
[0076] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.
[0077] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.
[0078] The Android Runtime consists of core libraries and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.
[0079] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.
[0080] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0081] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.
[0082] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.
[0083] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.
[0084] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0085] A 2D graphics engine is a graphics engine for 2D drawing.
[0086] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.
[0087] It should be noted that although the embodiments of this application are described using the Android system as an example, the basic principles are also applicable to electronic devices based on operating systems such as iOS or Windows.
[0088] Combination Figure 3 The software framework shown illustrates an example of how audio functionality is implemented during live streaming:
[0089] After the microphone captures the audio stream, it is transmitted to the media library through the kernel layer. The audio stream processed by the media library is transmitted to the live streaming API in the application framework layer, and then processed by the live streaming API before being transmitted to the live streaming application in the application layer.
[0090] Figure 4 An audio processing method disclosed in this application includes the following steps:
[0091] S401, In response to the selection command, display the coach model selection interface.
[0092] It is understood that selection commands can be triggered by user actions on the electronic device, including but not limited to: operations on the interactive interface and operations on buttons. Buttons can be physical or virtual. An example of such an operation is double-clicking the power button.
[0093] A coaching model is a combination of audio parameters that can achieve a certain audio effect. At least one coaching model can be pre-configured in an electronic device.
[0094] Understandably, the audio effect that users expect can be achieved by selecting a coach model.
[0095] In some implementations, the coach model selection interface includes a first interface, which is used to receive the attribute information of the coach model expected by the user. The attribute information of the coach model includes, but is not limited to: industry, gender, and style.
[0096] For example Figure 5a As shown, the industry options displayed on the first interface include: sales, training, hosting, and stand-up comedy. The gender options on the first interface include male and female. The style options on the first interface include character models, such as well-known live-streaming sales hosts, well-known teachers in the training industry, and well-known stand-up comedians.
[0097] Understandable, Figure 5a The controls shown are for illustrative purposes only, and this embodiment does not limit the style or operation of the controls.
[0098] The coach model selection interface also includes a second interface for receiving user attribute information. This user attribute information includes at least one of an audio sample and a text sample. The audio sample is a segment of the user's speech. The text sample is the text of a passage the user is about to record.
[0099] For example Figure 5b As shown, the second interface includes an audio sample input control "Please enter voice" and a text sample input control "Please import text". Users can click the audio sample input control to enter voice as an audio sample, and click the text sample input control to enter text. It is understood that when the user enters voice, no text is required; the electronic device obtains the text sample by performing text recognition on the voice.
[0100] Understandably, the style and operation of the controls on the second interface are not restricted.
[0101] As can be seen from the above explanation, the selection of the coaching model is based on two dimensions: the user's expected effect and the user's actual situation. Specifically, the attribute information of the coaching model received on the first interface represents the effect the user expects the audio to achieve, while the user's attribute information received on the second interface represents the actual situation of the audio the user is about to record. The combination of these two dimensions, serving as the selection criterion for the coaching model, satisfies the user's expectations while also being grounded in the user's actual conditions. Therefore, it can both achieve the expected live streaming effect and reflect the user's personalized needs.
[0102] Understandably, the display of the second interface can be triggered by a selection action on the first interface; for example, when a user selects an option on the first interface... Figure 5a After making a selection, a redirect will be triggered from the first screen to... Figure 5b The second interface shown. Alternatively, the first interface and the second interface can be displayed simultaneously in different areas of the electronic device's screen. In this case, the aforementioned triggering relationship does not exist between the first interface and the second interface.
[0103] S402. In response to the operation in the coach model selection interface, display the coach model recommendation interface.
[0104] Understandably, based on the information received from the first and second interfaces, at least one coaching model with a high degree of matching with the information received from the pre-configured coaching model library is selected as the recommended coaching model.
[0105] In some implementations, the method for selecting recommended coaching models based on the information received from the first interface and the second interface is as follows: candidate models that match the attribute information of the coaching models received from the first interface are selected, and then multiple recommended coaching models are selected from the candidate models according to the features extracted from the user's attribute information received from the second interface.
[0106] Understandably, multiple recommended coaching models can be sorted and displayed according to their matching degree.
[0107] It should be noted that if the coaching model determined by the information received from the first and second interfaces conflicts—for example, the industry option received by the first interface is "training," while the industry determined by the audio sample analysis received by the second interface is "stand-up comedy"—a third interface can be displayed. The third interface is used to display a conflict warning, indicating that the user's expected model conflicts with the model determined by the actual sample (at least one of the aforementioned audio and text samples). The third interface also receives secondary confirmation information, which is input by the user and selected from the conflicting information.
[0108] For example, such as Figure 5c As shown, the third interface displays the prompt message "Please reselect industry". After receiving the input information on the third interface, the recommended coaching model is filtered based on the input information and the information already obtained.
[0109] like Figure 5d As shown, the recommended coaching models are displayed in a list. The coaching model that appears earlier in the list matches the information received by the first and second interfaces to a higher degree.
[0110] Figure 5dIn the system, each coaching model can be displayed as an icon and a summary. The icon indicates gender information, and the summary includes a brief description of the voice, speaking speed, tone, and style. Optionally, the icon for each coaching model can be a control-style icon. Clicking the icon control plays the audio of that coaching model, allowing users to intuitively experience the audio style of that coaching model.
[0111] S403, In response to the operation on the coach model recommendation interface, confirm the coach model selected by the user.
[0112] For ease of distinction, the coaching model selected by the user will be referred to as the goal coaching model.
[0113] S404. After acquiring the audio stream, extract the speech pattern features of the audio stream.
[0114] In some implementations, an audio stream can be acquired by capturing the audio stream through a microphone.
[0115] In this embodiment, the speech pattern features include at least one of the following: voice features, speech rate features, and tone features. It is understood that the more audio features extracted, the more accurate the subsequent user reminders or audio processing will be.
[0116] In some implementations, speech features may include at least one technical parameter, such as pitch. The specific technical parameters included in the speech features can be pre-configured. Speech rate features and intonation features are similarly configured.
[0117] Audio features of an audio stream can be extracted using a pre-configured feature extraction model.
[0118] S405. Call the target coach model and the pre-configured speech pattern recognition model to obtain the speech pattern recognition results of the audio stream.
[0119] The speech pattern recognition model is composed of speech pattern features extracted from the audio segment that satisfies the user. The audio segment that satisfies the user can be pre-recorded and stored by the user, or it can be an audio segment extracted from a completed live stream recording. Examples of methods for obtaining the audio segment that satisfies the user can be found in S408-S409 and... Figure 7 As shown.
[0120] In some implementations, the way to obtain the speech pattern recognition results based on the above two models is as follows: if the number of speech pattern features in the speech pattern recognition model is not greater than the preset number, it means that the number of features in the speech pattern recognition model is insufficient. Therefore, in order to ensure the accuracy of the speech pattern recognition results, the target coach model is used to obtain the speech pattern recognition results of the audio stream. If the number of speech pattern features in the speech pattern recognition model is greater than the preset number, it means that the number of features in the speech pattern recognition model is sufficient to obtain accurate pattern recognition results. Therefore, the speech pattern recognition model is used to obtain the speech pattern recognition results of the audio stream.
[0121] The speech pattern recognition result is: speech pattern does not need adjustment or speech pattern needs adjustment. In response to the speech pattern recognition result indicating that speech pattern needs adjustment, execute at least one of S406 and S407:
[0122] S406. Display a prompt message indicating that the speaking style can be adjusted.
[0123] The speech mode adjustment prompts indicate how to adjust at least one of the following: speech tone, speech rate, and tone.
[0124] In some implementations, a prompt message indicating an adjustment to the speaking style is displayed, for example... Figure 6 As shown, the prompt message for adjusting the speaking style is "Please speak slower." After seeing this prompt message during the live stream, users can slow down their speech in subsequent live streams as instructed.
[0125] In other implementations, adjustment prompts are delivered via voice. Before the live stream, the user can select the desired display method through a settings interface. Understandably, the voice-based adjustment prompts are output through a specific channel to a designated receiving device worn by the streamer, while viewers will not hear the voice-based prompts.
[0126] S407. Call the target coach model to process the audio stream whose recognition result indicates that the speaking style needs to be adjusted, and output the processed audio stream.
[0127] The processing method can involve processing at least one of the speech, speech rate, and intonation. Taking speech as an example, at least one technical parameter of the speech features can be processed to achieve speech processing.
[0128] The selection criteria for S406 and S407 are as follows: They are executed according to a pre-configured method, allowing the user to determine the selected execution steps through user actions on the settings interface before the live broadcast. Alternatively, S406 can prompt for non-personalized parameters that are easily adjusted manually, such as speech rate, while S407 can handle more personalized parameters that are not easily adjusted manually, such as tone of voice. For ease of explanation, non-personalized parameters that are easily adjusted manually are referred to as Category 1 features, and more personalized parameters that are not easily adjusted manually are referred to as Category 2 features. Both Category 1 and Category 2 features can be pre-configured.
[0129] Furthermore, the first and second types of features can be adjusted based on the user's satisfaction with the processed audio stream, which will be explained in detail in conjunction with S409.
[0130] It is understandable that personalized parameters are those based on long-term speaking habits, making them difficult for users to adjust in a short period of time, i.e., not easy to adjust manually. Therefore, automatic adjustment helps to reduce the difficulty of operation for users.
[0131] Optionally, after the live stream ends, the following steps can be performed to optimize the speech recognition model using information from user feedback.
[0132] S408, in response to the live stream end command, displays the feedback interface.
[0133] The feedback interface is used to receive feedback information.
[0134] In some implementations, the feedback interface displays playback controls and evaluation controls for the audio frames output during the live stream. Understandably, users can listen to audio frames by using the playback controls and provide feedback on those frames using the evaluation controls. Figure 7 For example, the playback control for an audio frame is a play button, with the audio frame number displayed to the left of the play button. The evaluation control is a rating control, where the user inputs the rating by selecting the number of asterisks. This implementation focuses on collecting user feedback on the sound effects of the played audio. For this implementation, step S409 can be executed.
[0135] S409. Adjust the speech pattern recognition model based on the speech pattern characteristics of audio frames that meet the evaluation conditions.
[0136] In some implementations, the speech pattern features of audio frames that meet the evaluation conditions are incorporated into the speech pattern recognition model.
[0137] Specifically, if the evaluation information meets the conditions, the evaluation score can be greater than a preset score threshold. S409 enables the speech pattern recognition model to better meet user needs over time.
[0138] In other implementations, the first and second types of features are adjusted based on the evaluation information of the processed audio frames. Specifically, if the evaluation score of a processed audio feature is greater than a preset score threshold, it indicates that the user is satisfied with the automatic processing method, and this automatic processing method can be retained, with the parameter targeted by this automatic processing method being used as the second type of feature. Conversely, if the evaluation score of a processed audio feature is less than the preset score threshold, it indicates that the user is not satisfied with the automatic processing method, and this type of parameter is no longer processed automatically, with the parameter targeted by this automatic processing method being used as the first type of feature.
[0139] from Figure 4 As can be seen from the process shown, electronic devices have the function of adjusting or prompting the broadcaster to adjust the way of speaking reflected in the audio stream during the live broadcast, which is conducive to obtaining better live broadcast results.
[0140] Furthermore, the selection criteria for the coach model are based on the user's expectations and actual situation, and the speech pattern recognition model is composed of features extracted from audio segments that the user is satisfied with, thus making the adjusted or processed audio more in line with the user's needs.
[0141] Furthermore, the feedback information is used to adjust the speech recognition model, making the live audio better meet user expectations.
[0142] In summary, Figure 4 The audio processing method shown improves the audio function of electronic devices, making it easier to output audio that satisfies users.
[0143] Figure 4 The process shown is not limited to live streaming scenarios; it can also be applied to practice scenarios, such as practice sessions in preparation for a live stream. In contrast to live streaming, in practice scenarios, electronic devices only capture the video stream without transmitting it externally.
[0144] The audio processing methods in practice scenarios, and Figure 4 Compared to the process shown, the difference lies in the addition of a step to display the practice text, and the adjustment of the display method of prompts to suit the speaking style in the practice scenario.
[0145] Because live streaming requires high real-time performance, therefore, Figure 4 In the process shown, the display of adjustment prompts can be combined with automatic processing. However, in practice scenarios, the automatic processing step S407 can be omitted.
[0146] Furthermore, in live streaming scenarios, to avoid distracting users and affecting the streaming experience, the prompts should be kept as brief as possible. In practice scenarios, however, the prompts can be more detailed to ensure users achieve sufficient practice results.
[0147] Another audio processing method disclosed in the embodiments of this application, such as... Figure 8 As shown:
[0148] S801-S803 are similar to S401-S403, the difference being:
[0149] Users can trigger the practice mode through the practice control. In practice mode, they can follow... Figure 5a- Figure 5c The target coaching model is selected through a method where the text sample entered in the second interface can be a practice text, such as a product description or training courseware.
[0150] S804 introduces a new step: After starting audio streaming in practice mode, the annotated practice text is displayed on the screen of the electronic device. Annotations include, but are not limited to, stress marks and pause marks. For example, text requiring stress is highlighted in bold in the practice text, and pause marks are added within sentences. In some implementations, a pre-trained model can be used to determine the annotations in the text.
[0151] Understandably, users can transmit audio based on the annotated exercise text.
[0152] S805-S806 are the same as S404-S405, so they will not be described again here.
[0153] In S807, adjusting the prompt message can be done in addition to... Figure 4 In addition to the adjustment prompts shown in the flowchart, you can also add error messages compared to the annotated exercise text. For example, displaying accents and pauses that are inconsistent with those in the annotated exercise text using specific colors, shapes, etc.
[0154] S808-S809 are the same as S408-S409, so they will not be repeated here. It is understandable that the features of audio frames whose evaluation information meets the conditions in S808 can be added not only to the speech pattern recognition model in practice mode but also to the speech pattern recognition model in live streaming mode. The goal is to quickly make the speech pattern recognition model more aligned with user needs.
[0155] Understandably, in Figure 4 or Figure 8 In the illustrated process, the steps for selecting the training model, namely S401-S403 and S801-S803, are optional and can be omitted. In this case, only the speech pattern recognition model is used to obtain the speech pattern recognition result of the audio stream. Alternatively, these steps can be performed once. After the user selects a coach model, the already selected coach model will be used as long as the user does not trigger the selection of a new coach model.
Claims
1. An audio processing method, applied to an electronic device, characterized in that, include: Get the audio stream; The model is invoked to obtain the speech pattern recognition result of the audio stream based on the speech pattern characteristics of the audio stream. The model includes a target coach model, which is a model configured based on at least one of the audio input party's expected audio effect and the audio input party's attribute information. The audio input party's expected audio effect is represented based on the attribute information of the target coach model, which includes: industry, gender, and style. The speech pattern features are divided into a first type of features and a second type of features, where the second type of features are less easily adjusted manually than the first type of features. In response to the speech pattern recognition result of the audio stream indicating that the speech pattern needs to be adjusted, the system displays adjustment prompt information for the first type of feature, and / or outputs the audio stream adjusted for the second type of feature.
2. The method according to claim 1, characterized in that, The attribute information of the audio input includes: text sample.
3. The method according to claim 1 or 2, characterized in that, The calling model obtains the speech pattern recognition result of the audio stream based on the speech pattern characteristics of the audio stream, including: The target coach model and the speech pattern recognition model are invoked, and the speech pattern recognition result of the audio stream is obtained based on the speech pattern features of the audio stream. The speech pattern features include at least one of the speech features, intonation features, and speech rate features of the audio stream.
4. The method according to claim 1 or 2, characterized in that, The configuration process for the target coaching model includes: In response to a selection command, a coach model selection interface is displayed, which is used to receive at least one of the audio input party's expected audio effect and the audio input party's attribute information. In response to the operation in the coach model selection interface, the coach model recommendation interface is displayed; In response to the operation in the coach model recommendation interface, the selected target coach model is confirmed.
5. The method according to claim 4, characterized in that, The display of the coach model selection interface includes: Display the first interface, and / or display the second interface; The first interface includes: industry options, gender options, and character model options; The second interface includes a receiving control for at least one of audio samples and text samples.
6. The method according to claim 5, characterized in that, Before displaying the coach model recommendation interface in response to an operation in the coach model selection interface, the method further includes: In response to a conflict between the audio input device's expected audio effect and the audio input device's attribute information, a third interface is displayed, which is used to receive secondary confirmation information regarding the conflicting information.
7. The method according to claim 4, characterized in that, The interface for displaying coach model recommendations includes: The coach model recommendation interface displays information including recommended coach models; The information of the first recommended coach model includes: an icon control and summary information. The icon control is used to trigger the playback of the audio of the first recommended coach model. The summary information includes: an introduction to at least one of the categories and styles of the voice, speech rate, and tone of the first recommended coach model. The first recommended coach model can be any recommended coach model.
8. The method according to claim 1 or 2, characterized in that, Also includes: In response to an audio acquisition end command, a feedback interface is displayed. The feedback interface is used to receive feedback information, which includes evaluation information for each audio frame in the output audio stream.
9. The method according to claim 8, characterized in that, Also includes: Based on the speech pattern characteristics of audio frames that meet the evaluation information conditions, the speech pattern recognition model is adjusted.
10. An electronic device, characterized in that, include: Display screen, processor, and memory; The memory is used to store an application program, and the processor is used to run the application program to implement the audio processing method according to any one of claims 1-9.
11. A readable storage medium having an application program stored thereon, characterized in that, When the application is run on a computer device, the audio processing method according to any one of claims 1-9 is implemented.
Citation Information
Patent Citations
Method and device for hinting friendlily
CN103269405A
Video dubbing method and device based on voice synthesis, computer equipment and medium
CN111031386A