Audio processing method and related device

By outputting information about the degree of data density in multiple periods of audio data and generating virtual scene videos, the problem of users obtaining inaccurate information from converted text is solved, and user experience and information positioning efficiency is improved.

CN120020701APending Publication Date: 2025-05-20HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311545575.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-17
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

In the prior art, the information obtained by users from the converted text is inaccurate, resulting in a poor user experience.

Method used

By acquiring the audio data to be processed, information indicating the data density of the audio data in each period of multiple periods is output, allowing the user to quickly locate periods with higher or lower data density and generate virtual scene video based on the audio scene.

Benefits of technology

It improves the efficiency of users to quickly locate target information in audio data, improves the user experience, and ensures the accuracy of information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120020701A_ABST
    Figure CN120020701A_ABST
Patent Text Reader

Abstract

The invention provides an audio processing method and a related device, which can be applied to the technical field of terminals. In the technical scheme provided by the invention, after the terminal device obtains the to-be-processed audio data, the first information can be output based on the to-be-processed audio data, and the first information can indicate the data density degree of the to-be-processed audio data in each of a plurality of time periods. In the method, the user can quickly position the time period containing more information based on the data density of each time period in the plurality of time periods, so that the efficiency of acquiring the target information by the user is improved, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of terminals, and particularly to an audio processing method and related devices. Background Art

[0002] Currently, audio data can be obtained in a terminal device and audio can be played based on the audio data. When a user uses the terminal device to play audio, there is usually a need to quickly and accurately locate and find target information.

[0003] In the prior art, audio data can be converted into text to help the user quickly browse and find target information. However, during the process of converting audio data into text, it may happen that the audio data cannot be recognized, resulting in inaccurate information obtained by the user from the converted text and poor user experience. Summary of the Invention

[0004] This application provides an audio processing method and related devices, which are used to solve the problem that the information obtained by the user from the converted text in the prior art is inaccurate.

[0005] In a first aspect, this application provides an audio processing method, and the method includes: obtaining audio data to be processed; outputting first information based on the audio data to be processed, where the first information indicates the data density of the audio data to be processed in each of multiple time periods.

[0006] As an example, the audio data to be processed can be audio data obtained by recording.

[0007] These multiple time periods can be preset in advance, and the duration of each of these multiple time periods can also be preset in advance. For example, the start time and end time of each time period can be preset in advance.

[0008] Optionally, the duration of each of these multiple time periods can be equal.

[0009] Optionally, the terminal device can first determine these multiple time periods, then parse the audio data to be processed to obtain the data density of the audio data to be processed in each of these multiple time periods, and then output the first information.

[0010] In this method, the data density can represent the amount of data per unit duration. When the duration of each of these multiple time periods is equal, the data density of the audio data to be processed in each time period can be characterized by the amount of audio data of the audio data to be processed in each time period.

[0011] In this method, based on the first information, the user can determine the time periods with a higher data density and the time periods with a lower data density in the audio data to be processed. Or rather, based on the first information, the user can determine the time periods with a larger amount of data and the time periods with a smaller amount of data in the audio data to be processed. Generally, the time periods with a larger amount of data contain more information. In this way, the user can quickly locate the time periods with more information in the audio data to be processed, which is beneficial to improving the efficiency of the user in obtaining the target information and enhancing the user experience.

[0012] In some possible implementation manners, the method further includes: obtaining second information, where the second information indicates the audio scene of the audio data to be processed; generating a first video based on the second information, where the first video includes a virtual scene corresponding to the audio scene of the audio data to be processed.

[0013] Taking the audio data obtained by recording as the audio data to be processed as an example, the audio scene of the audio data to be processed can be the recording scene, or rather, the real scene during recording.

[0014] In this example, the virtual scene corresponding to the audio scene of the audio data to be processed can be a virtual scene that simulates the real scene during recording. The content included in the virtual scene can be the same as or different from the content in the real scene. For example, as long as the scene types of the audio scene and the virtual scene are the same, such as both the audio scene and the virtual scene being a meeting room, etc.

[0015] In this method, the user can intuitively and accurately obtain the target information based on the generated first video, which is beneficial to improving the user experience.

[0016] In some possible implementation manners, the audio data to be processed includes the audio data of one or more speakers, and the virtual scene includes virtual characters corresponding to at least one of the one or more speakers.

[0017] Among them, the virtual characters corresponding to each of the at least one speaker can be virtual characters that simulate each speaker.

[0018] In this method, the terminal device can determine, through voice recognition technology, that the audio data to be processed includes the audio data of one or more speakers, then determine, through sound source localization technology, the position of at least one of the one or more speakers in the audio scene, and then determine, based on the position of each of the at least one speaker in the audio scene, the position of the virtual character corresponding to each speaker in the virtual scene.

[0019] In this method, the terminal device can simulate a speaker through a virtual character and load the virtual character into a virtual scene, so that the generated virtual scene can approximate the actual audio scene, which is beneficial to improving the user experience.

[0020] In some possible implementation manners, the positions of the virtual characters corresponding to each speaker among the at least one speaker in the virtual scene are the same as the positions of each speaker in the audio scene.

[0021] In this method, the positions of the virtual characters corresponding to each speaker among the at least one speaker in the virtual scene are the same as the positions of each speaker in the audio scene, so that the generated virtual scene is closer to the audio scene, which is beneficial to improving the user experience.

[0022] In some possible implementation manners, the method further includes: playing the first video, where the picture focus of the first video in the first period among the multiple periods is on the virtual character corresponding to the first speaker among the one or more speakers, and the first period is the period when the first speaker is speaking.

[0023] That the picture focus of the first video in the first period is on the virtual character corresponding to the first speaker can be understood as: in each video frame of the first video located in the first period, the proportion of the picture of the virtual character of the first speaker in each video frame is greater than the proportion of the pictures of the virtual characters of other speakers except the first speaker in each video frame.

[0024] In this method, the terminal device can enlarge the proportion of the picture of the speaker in the first video, which can provide a more realistic audio-visual experience for the user and is beneficial to improving the user experience.

[0025] In some possible implementation manners, the first video displays the action information of the virtual character corresponding to the first speaker in the first period, and the action information of the virtual character corresponding to the first speaker includes the mouth action of the virtual character corresponding to the first speaker.

[0026] In this method, the first video may include the action information of the virtual character corresponding to the speaker, which can provide a more realistic audio-visual experience for the user and is beneficial to improving the user experience.

[0027] In some possible implementation manners, the method further includes: obtaining third information, where the third information indicates that the user clicks on the virtual character corresponding to the second speaker among the one or more speakers; amplifying the voice of the second speaker.

[0028] In this method, the terminal device can amplify the voice of the second speaker based on the third information, which is beneficial for the user to accurately obtain the target information required by the user from the amplified voice.

[0029] In some possible implementation manners, the picture focus of the first video in the second period of the multiple periods is on the virtual character corresponding to the second speaker among one or more speakers. The start time of the second period is the time when the user clicks on the virtual character corresponding to the second speaker, and the end time of the second period is the time when the user clicks on other speakers except the second speaker among the one or more speakers.

[0030] In this method, the terminal device can also amplify the screen occupation ratio of the second speaker in the first video, which can provide a more realistic audio-visual experience for the user and is beneficial to improving the user experience.

[0031] In some possible implementation manners, the data density of each period among the multiple periods is displayed in one or more of the following ways: energy bar, color.

[0032] As an example, when the data density is displayed by an energy bar, the more energy bars there are in the period with a higher data density. Or rather, the more energy bars there are in the period with a larger data volume.

[0033] As Figure 4 shown, assuming that these multiple periods include period 1 and period 2, the durations of period 1 and period 2 are equal, and the data volume of the audio data to be processed in period 1 is greater than the data volume of the audio data to be processed in period 2, then the energy bar of period 1 is more than the energy bar of period 2. Or rather, the length of the energy bar of period 1 is greater than the length of the energy bar of period 2. In this example, the shaded part below the progress bar of each period can represent the energy bar of each period. In this example, the data density of each period among the multiple periods is displayed below the audio progress bar.

[0034] Optionally, the energy bars of different periods can also be as Figure 5 shown. In this example, the data density of each period among the multiple periods can be represented by the density of the energy bar, and the more dense the energy bar of the period with a higher data density.

[0035] Assuming that these multiple periods include period 1 and period 2, the durations of period 1 and period 2 are equal, and the data volume of the audio data to be processed in period 1 is greater than the data volume of the audio data to be processed in period 2, then the energy bar of period 1 is denser than the energy bar of period 2.

[0036] As another example, when the data density is displayed by color, the colors of different time periods with different data densities can be different. For example, the color of a time period with a higher data density is darker.

[0037] As Figure 6 shown, assume that among these multiple time periods, there are time period 1 and time period 2, and the durations of time period 1 and time period 2 are equal. If the amount of audio data to be processed in time period 1 is greater than the amount of audio data to be processed in time period 2, then the color of the position of time period 1 in the progress bar is different from the color of the position of time period 2 in the progress bar.

[0038] Regarding the color in the progress bar, the blank part can represent the color of time period 2 in the progress bar, and the color of the position of time period 1 in the progress bar is darker than the color of the position of time period 2 in the progress bar.

[0039] Optionally, when the durations of each of these multiple time periods are equal, the data density of each of these multiple time periods can also be represented by the amount of data in each time period.

[0040] As an example, the amount of data in each time period can be represented by the bit rate or bit value of the audio data in each time period.

[0041] As Figure 7 shown, assume that these multiple time periods include time period 1, time period 2, time period 3, and time period 4, and the amount of data in each of these four time periods is represented by the bit rate of the audio data in each time period.

[0042] In some possible implementation manners, the method further includes: determining, from the multiple time periods, the time periods in which the amount of audio data exceeds a preset data amount threshold; marking the time periods in which the amount of audio data exceeds the preset data amount threshold.

[0043] In this method, by marking the time periods in which the amount of audio data exceeds the preset data amount threshold, the user can determine the time periods in which the amount of audio data exceeds the preset data amount threshold in the audio or video, and then obtain the target information from the time periods in which the amount of audio data exceeds the preset data amount threshold, which is beneficial for the user to quickly obtain the target information.

[0044] In some possible implementation manners, the method further includes: performing semantic analysis on the audio data in the time periods in which the amount of audio data exceeds the preset data amount threshold to obtain key information in the time periods in which the amount of audio data exceeds the preset data amount threshold; outputting the key information.

[0045] In this method, by displaying key information, the user can determine a period containing the target information required by the user from at least one period in which the audio data volume exceeds a preset data volume threshold, and then obtain the target information from the audio or video in the period containing the target information required by the user, which is beneficial for the user to quickly obtain the target information.

[0046] In a second aspect, the present application provides an audio processing device, including modules or units for implementing the methods in the first aspect and any possible implementation manner of the first aspect. It should be understood that each module or unit can implement corresponding functions by executing a computer program.

[0047] As an example, the audio processing device may include an acquisition module and a processing module.

[0048] Among them, the acquisition module can be used to acquire audio data to be processed.

[0049] The processing module can be used to output first information based on the audio data to be processed, and the first information indicates the data density of the audio data to be processed in each period of multiple periods.

[0050] In some possible implementation manners, the acquisition module can also be used to acquire second information, and the second information indicates the audio scene of the audio data to be processed.

[0051] The processing module can also be used to generate a first video based on the second information, and the first video includes a virtual scene corresponding to the audio scene of the audio data to be processed.

[0052] In some possible implementation manners, the audio data to be processed includes audio data of one or more speakers, and the virtual scene includes virtual characters corresponding to at least one of the one or more speakers.

[0053] In some possible implementation manners, the position of the virtual character corresponding to each of the at least one speaker in the virtual scene is consistent with the position of each speaker in the audio scene.

[0054] In some possible implementation manners, the processing module can also be used to play the first video, and the picture focus of the first video in the first period of the multiple periods is on the virtual character corresponding to the first speaker among one or more speakers, and the first period is the period when the first speaker is speaking.

[0055] In some possible implementations, the first video displays the action information of the virtual character corresponding to the first speaker during the first time period, and the action information of the virtual character corresponding to the first speaker includes the mouth actions of the virtual character corresponding to the first speaker.

[0056] In some possible implementations, the obtaining module may also be used to obtain third information, where the third information indicates that the user clicks on the virtual character corresponding to the second speaker among the one or more speakers.

[0057] The processing module may also be used to amplify the voice of the second speaker.

[0058] In some possible implementations, the picture focus of the first video during the second time period of the multiple time periods is on the virtual character corresponding to the second speaker among the one or more speakers. The start time of the second time period is the time when the user clicks on the virtual character corresponding to the second speaker, and the end time of the second time period is the time when the user clicks on other speakers among the one or more speakers except the second speaker.

[0059] In some possible implementations, the data density of each time period among the multiple time periods is displayed in one or more of the following ways: energy bar, color.

[0060] In some possible implementations, the processing module may also be used to determine, from the multiple time periods, the time periods in which the audio data volume exceeds a preset data volume threshold.

[0061] The processing module may also be used to mark the time periods in which the audio data volume exceeds the preset data volume threshold.

[0062] In some possible implementations, the processing module may also be used to perform semantic analysis on the audio data within the time periods in which the audio data volume exceeds the preset data volume threshold, to obtain the key information within the time periods in which the audio data volume exceeds the preset data volume threshold.

[0063] The processing module may also be used to output the key information.

[0064] In a third aspect, the present application provides an audio processing device, including a processor, where the processor is used to execute the method according to any one of the possible implementations in the first aspect.

[0065] The device may further include a memory for storing instructions and data. The memory is coupled to the processor, and when the processor executes the instructions stored in the memory, the methods described in the above aspects can be implemented. The device may further include a communication interface for the device to communicate with other devices. Exemplarily, the communication interface may be a transceiver, a circuit, a bus, a module, or other types of communication interfaces.

[0066] In a fourth aspect, the present application provides a computer-readable storage medium storing program code for a device to execute, the program code including instructions for implementing the method described in any possible implementation manner of the first aspect.

[0067] In a fifth aspect, the present application provides a computer program product including instructions, which when the computer program product runs on a device, enables the device to implement the method described in any possible implementation manner of the first aspect.

[0068] It can be understood that the effects obtainable in the second aspect to the fifth aspect can be referred to the description in the first aspect, and will not be elaborated herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 It is a schematic structural diagram of a terminal device provided by an embodiment of the present application;

[0070] Figure 2 It is a software structure block diagram of a terminal device applicable to an embodiment of the present application;

[0071] Figure 3 It is a schematic flowchart of an audio processing method provided by the present application;

[0072] Figure 4 It is a schematic diagram of a display manner of first information provided by an embodiment of the present application;

[0073] Figure 5 It is a schematic diagram of a display manner of first information provided by another embodiment of the present application;

[0074] Figure 6 It is a schematic diagram of a display manner of first information provided by yet another embodiment of the present application;

[0075] Figure 7 It is a schematic diagram of a display manner of first information provided by yet another embodiment of the present application;

[0076] Figure 8 It is a schematic diagram of a first user interface provided by an embodiment of the present application;

[0077] Figure 9 It is a schematic diagram of a second user interface provided by an embodiment of the present application;

[0078] Figure 10 Schematic diagram of the first user interface provided for another embodiment of this application;

[0079] Figure 11 Schematic diagram of the images of the first speaker at different moments in the video provided for an embodiment of this application;

[0080] Figure 12 Schematic diagram of a video image provided for an embodiment of this application;

[0081] Figure 13 Schematic diagram of the first video provided for an embodiment of this application;

[0082] Figure 14 Schematic diagram of the first video provided for another embodiment of this application;

[0083] Figure 15 Schematic diagram of the audio processing device provided for an embodiment of this application;

[0084] Figure 16 Schematic diagram of the audio processing device provided for another embodiment of this application. Detailed implementation manners

[0085] Next, the technical solutions in the embodiments of this application will be described in conjunction with the accompanying drawings in the embodiments of this application.

[0086] For the convenience of clearly describing the technical solutions in the embodiments of this application, in the embodiments of this application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. For example, the first information and the second information are only used to distinguish different information, and no limitation is imposed on their sequence. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and "first", "second", etc. do not necessarily mean different.

[0087] In the embodiments of this application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c can represent: a, b, c; a and b; a and c; b and c; or a, b, and c. Where a, b, and c can be single or multiple.

[0088] The method of the present application can be applied to a terminal device, which can be a mobile phone, a tablet computer, a personal computer (PC), a smart screen, a vehicle-mounted device, and wearable terminal devices such as smart watches, and can also be various teaching aids (such as learning machines, early education machines), smart toys, portable robots, personal digital assistants (PDAs), augmented reality (AR) devices, virtual reality (VR) devices, etc., or devices with mobile office functions, devices with smart home functions, devices with audio-visual entertainment functions, devices supporting smart travel, etc. It should be understood that the specific technologies and specific device forms adopted by the terminal device in the embodiments of the present application are not limited.

[0089] In order to better understand the embodiments of the present application, the hardware structure of the terminal device in the embodiments of the present application will be introduced below. Exemplarily, Figure 1 FIG. is a schematic structural diagram of a terminal device provided by an embodiment of the present application.

[0090] The terminal device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.

[0091] Optionally, the above sensor module 180 may include a pressure sensor 180A and a touch sensor 180B, etc.

[0092] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the terminal device. In other embodiments of the present application, the terminal device may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0093] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, etc. Among them, different processing units may be independent devices or integrated in one or more processors. A memory may also be provided in the processor 110 for storing instructions and data.

[0094] The USB interface 130 is an interface that complies with the USB standard specification. Specifically, it may be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 130 can be used to connect a charger to charge the terminal device, and can also be used to transfer data between the terminal device and peripheral devices. It can also be used to connect headphones to play audio through the headphones. This interface can also be used to connect other terminal devices, such as AR devices, etc.

[0095] The charging management module 140 is used to receive a charging input from a charger. Among them, the charger may be a wireless charger or a wired charger. While charging the battery 142, the charging management module 140 can also charge the terminal device through the power management module 141.

[0096] The power management module 141 is used to connect the battery 142, the charging management module 140 and the processor 110. The power management module 141 receives inputs from the battery 142 and / or the charging management module 140 and supplies power to the processor 110, the internal memory 121, the external memory, the camera 193, the display screen 194, the audio module 170, the wireless communication module 160, etc. In some embodiments, the power management module 141 may also be provided in the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may also be provided in the same device.

[0097] Optionally, in some embodiments, the power management module 141 may also be referred to as a power management unit (PMU). The PMU can be used to obtain the status information of the battery 142, such as the remaining power and temperature of the battery 142, etc.

[0098] The wireless communication function of the terminal device can be implemented through antenna 1, antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor, and the baseband processor, etc.

[0099] Antenna 1 and Antenna 2 are used for transmitting and receiving electromagnetic wave signals. The antennas in the terminal device can be used to cover single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas.

[0100] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc. applied to the terminal device. The mobile communication module 150 can include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves by Antenna 1, filter and amplify the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves through Antenna 1 and radiate it out.

[0101] The wireless communication module 160 can provide solutions for wireless communications including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), etc. applied to the terminal device. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves through Antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signals to be transmitted from the processor 110, perform frequency modulation and amplification on them, and convert them into electromagnetic waves through Antenna 2 and radiate them out.

[0102] In some embodiments, Antenna 1 of the terminal device is coupled to the mobile communication module 150, and Antenna 2 is coupled to the wireless communication module 160, so that the terminal device can communicate with the network and other devices through wireless communication technologies.

[0103] The camera 193 is used to capture static images or videos. In some embodiments, the terminal device can include one or N cameras 193, where N is a positive integer greater than 1.

[0104] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. In some embodiments, the terminal device can include one or N display screens 194, where N is a positive integer greater than 1.

[0105] In an embodiment of the present application, on the display screen 194, the interfaces of preset applications, the interface of the volume bar, etc. can be displayed. Among them, on the interface of the preset application, there can be included: video frame images, video duration, etc. When the video frame image is displayed on the display screen 194, the volume bar displayed on the display screen 194 can indicate the current volume value.

[0106] In some embodiments of the present application, the preset application can include but is not limited to application programs that can play audio or video, such as video playback applications, audio playback applications, game applications, shopping applications, and social applications, etc.

[0107] The terminal device can implement the display function through a graphics processing unit (GPU), the display screen 194, the application processor, etc. The GPU is a microprocessor for image processing, and is connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The GPU can also be called a display core, a visual processor, a display chip, etc.

[0108] The terminal device can implement the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, the application processor, etc.

[0109] The video codec is used to compress or decompress digital videos. The terminal device can support one or more video codecs. In this way, the terminal device can play or record videos in multiple coding formats, such as: Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc. In addition, the terminal device can implement the encapsulation and playback of audio and video data and screen recording data, etc.

[0110] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the terminal device. The external memory card communicates with the processor 110 through the external memory interface 120 to implement the data storage function. For example, save files such as music and videos in the external memory card.

[0111] The internal memory 121 can be used to store computer-executable program codes, and the executable program codes include instructions. The internal memory 121 can include a program storage area and a data storage area. Among them, the program storage area can store application programs required for at least one function in the operating system (such as the sound playback function, the image playback function, etc.). The data storage area can store data created during the use of the terminal device (such as audio data, phone book, etc.).

[0112] The terminal device can implement audio functions through the audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and the application processor, etc. For example, music playback, recording, etc.

[0113] The audio module 170 is used to convert digital audio information into an analog audio signal for output, and is also used to convert analog audio input into digital audio signals. The speaker 170A, also known as the "loudspeaker", is used to convert an audio electrical signal into a sound signal. The terminal device can listen to music or hands-free calls through the speaker 170A. The receiver 170B, also known as the "earpiece", is used to convert an audio electrical signal into a sound signal. When the terminal device answers a call or a voice message, the voice can be listened to by holding the receiver 170B close to the ear. The microphone 170C, also known as the "microphone" or "transmitter", is used to convert a sound signal into an electrical signal. The headphone jack 170D is used to connect a wired headphone.

[0114] The pressure sensor 180A is used to sense pressure signals and can convert the pressure signals into electrical signals. In some embodiments, the pressure sensor 180A can be disposed on the display screen 194. When a touch operation acts on the display screen 194, the terminal device can detect the intensity of the touch operation according to the pressure sensor 180A. The terminal device can also calculate the position of the touch according to the detection signal of the pressure sensor 180A. The touch sensor 180B can be disposed on the display screen 194, and the touch sensor 180B and the display screen 194 form a touch screen, also known as a "touch control screen". The touch sensor 180B is used to detect touch operations acting thereon or nearby. The touch sensor 180B can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In other embodiments, the touch sensor 180B can also be disposed on the surface of the terminal device, at a different position from that of the display screen 194.

[0115] The keys 190 include a power-on key, volume keys, etc. The keys 190 can be mechanical keys or touch keys. The terminal device can receive key inputs and generate key signal inputs related to the user settings and function controls of the terminal device. The motor 391 can generate a vibration prompt. The indicator 192 can be an indicator light, which can be used to indicate the charging state, power change, and can also be used to indicate messages, missed calls, notifications, etc. The timer 196 can be used to record time information. For example, when the terminal device records a video, it can record the start time and end time of the video, and can also record the duration of the recorded video.

[0116] In an embodiment of the present application, by way of example, the software system of the terminal device may adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. The layered architecture may adopt the Android system, the iOS system, or other operating systems, which is not limited in the embodiments of the present application. The following takes Figure 2 as an example to exemplarily illustrate the software structure of the terminal device.

[0117] Figure 2 FIG. is a block diagram of the software structure of a terminal device applicable to the embodiments of the present application. The layered architecture divides the software system of the terminal device into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the system may include an applications layer, an application framework layer, an Android runtime, system libraries, a hardware abstraction layer (HAL), and a kernel layer.

[0118] The applications layer may include a series of application packages, and the applications layer runs the applications by calling the application programming interfaces (APIs) provided by the application framework layer.

[0119] For example, the application packages may include applications such as a camera, a gallery, a calendar, a call, a map, a navigation, a wireless local area network (WLAN), Bluetooth, music, video, short message, and recording. Of course, the applications layer may also include third-party application packages, such as social applications, third-party music applications, third-party video applications, payment applications, shopping applications, bank applications, chat applications, or financial management applications, etc., which is not limited in the present application.

[0120] Among them, the social application may have the function of playing video and audio. In response to the user's operation of playing the video, the terminal device may play the video frame and the corresponding audio. In a possible design of the present application, the social application may, in response to the user's operation of playing the video, call the media playback interface to start the system's (one or more) players.

[0121] In some embodiments of the present application, during the process of the terminal device playing audio, the terminal device may set the volume level of the audio in response to the user's operation. In the case where the terminal device automatically plays a video, the terminal device may play the audio corresponding to the video based on the set volume level.

[0122] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions. For example, it may include an activity manager, a window manager, a content provider, a view system, a resource manager, and an audio service, etc. The embodiments of the present application do not impose any restrictions on this.

[0123] Among them, the audio service can be started during the boot stage of the terminal device, and the audio service can be used to transfer and save the relevant audio data information of the audio module.

[0124] The Android runtime includes a core library, a virtual machine, and a virtual machine monitor. The Android runtime is responsible for the scheduling and management of the Android system. The core library consists of two parts: one part is the functional functions that need to be called by the Java language, and the other part is the core library of Android. The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0125] The system library can include modules with multiple functions, such as: a surface manager, a media library, a 3D graphics processing library, and an identification algorithm module, etc.

[0126] The surface manager is used to manage the display subsystem and provides the fusion of 2D layers and 3D layers for multiple applications. The media library supports the playback and recording of multiple common audio and video formats, as well as static image files, etc. The media library can support multiple audio and video coding formats, such as: JPG, PNG, etc. The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc. The identification algorithm module can be used for sign language recognition, speech recognition, and text semantic recognition. Among them, sign language recognition refers to recognizing speech or text as sign language, speech recognition refers to recognizing sign language or text as speech, and text semantic recognition refers to recognizing sign language or speech as text.

[0127] The HAL layer is an encapsulation of the Linux kernel driver, providing an interface upward and shielding the implementation details of the low-level hardware.

[0128] The HAL layer may include a video HAL, an audio HAL, and a camera HAL, etc.

[0129] The kernel layer is the layer between the hardware and the software. The kernel layer is used to drive the hardware so that the hardware works. The kernel layer at least includes a display driver, a camera driver, an audio driver, a sensor driver, etc. Among them, the audio driver is the driver layer for audio-related devices (such as speakers), and is mainly responsible for the interaction with the hardware.

[0130] In the embodiments of the present application, the display driver may include the driver corresponding to the display, and the audio driver may include the driver corresponding to the speaker, the driver corresponding to the receiver, and the driver corresponding to the microphone.

[0131] The hardware layer includes a display, a camera, a speaker, a receiver, and a microphone, etc.

[0132] It can be understood that Figure 2 In the illustrated embodiment, the software structure of the terminal device is only a simple example and does not limit the scope of the present application.

[0133] In the present application, the terminal device can obtain audio data and play audio based on the audio data. An example of the audio is recorded audio, and correspondingly, the audio data can be the audio data obtained by recording. Subsequently, the present application will take the recorded audio as an example for illustration.

[0134] In one case, the recorded audio can be recorded by the terminal device. For example, the user can input a recording instruction on the terminal device, and the terminal device can record based on the recording instruction, and the recorded audio data can be obtained after the recording is completed.

[0135] In another case, the terminal device can also obtain the recorded audio from other devices, and the architecture of the other devices can also be as Figure 1 and Figure 2 shown. For example, the user can input a recording instruction on the other device, and the other device can record based on the recording instruction, and the recorded audio data can be obtained after the recording is completed. After the other device obtains the recorded audio data, it can send the recorded audio data to the terminal device. Correspondingly, the terminal device can receive the recorded audio data.

[0136] When the user uses the terminal device to play audio, there is usually a need to quickly and accurately locate and find the target information, and the target information can be the information required by the user.

[0137] In the prior art, the audio data can be converted into text to help the user quickly browse and find the target information. However, in the process of converting the audio data into text, there may be a situation where some audio data cannot be converted into text because the audio data cannot be recognized, or there may be a situation where the converted text is inconsistent with the information of the audio data, resulting in inaccurate information obtained by the user from the converted text and poor user experience.

[0138] To this end, the present application provides an audio processing method to solve the problem that the information obtained by the user from the converted text in the prior art is inaccurate.

[0139] In the technical solution of the present application, after the terminal device obtains the audio data to be processed, it can output a first piece of information based on the audio data to be processed, and the first piece of information can indicate the data density of the audio data to be processed in each of multiple time periods.

[0140] Among them, these multiple time periods can be preset in advance, and the duration of each of these multiple time periods can also be preset in advance. For example, the start time and end time of each time period can be preset in advance.

[0141] Optionally, the duration of each of these multiple time periods can be equal.

[0142] Optionally, the terminal device can first determine these multiple time periods, then parse the audio data to be processed to obtain the data density of the audio data to be processed in each of these multiple time periods, and then output the first piece of information.

[0143] In this method, the data density can represent the amount of data per unit time. When the duration of each of these multiple time periods is equal, the data density of the audio data to be processed in each time period can be characterized by the amount of audio data of the audio data to be processed in each time period.

[0144] As an example, assume that these multiple time periods can include a first time period and a second time period, the duration of the first time period is equal to the duration of the second time period, the data density of the audio data to be processed in the first time period can be characterized by the amount of data of the audio data to be processed in the first time period, and the data density of the audio data to be processed in the second time period can be characterized by the amount of data of the audio data to be processed in the second time period.

[0145] Assume that the duration of the first time period is equal to the duration of the second time period. If the amount of data of the audio data to be processed in the first time period is greater than the amount of data of the audio data to be processed in the second time period, it means that the data density of the audio data to be processed in the first time period is higher than the data density of the audio data to be processed in the second time period, that is, the audio data of the audio data to be processed in the first time period is denser than the audio data of the audio data to be processed in the second time period.

[0146] In this method, based on the first information, the user can determine the time periods with a higher data density and the time periods with a lower data density in the audio data to be processed. Or rather, based on the first information, the user can determine the time periods with a larger amount of data and the time periods with a smaller amount of data in the audio data to be processed. Generally, the time periods with a larger amount of data contain more information. In this way, the user can quickly locate the time periods with more information in the audio data to be processed, which is beneficial to improving the efficiency of the user to obtain the target information and enhancing the user experience.

[0147] Furthermore, the terminal device can also generate a first video, which can include a virtual scene corresponding to the audio scene of the audio data to be processed.

[0148] Optionally, the audio data to be processed can include the audio data of one or more speakers, and the first video can include virtual characters corresponding to at least one of the one or more speakers.

[0149] Optionally, the position of each speaker in the virtual scene can be the same as the position of each speaker in the audio scene.

[0150] In this method, the user can intuitively and accurately obtain the target information based on the generated first video, which is beneficial to improving the user experience.

[0151] Next, this application will be combined with Figures 3 to 16 , and the method of this application will be introduced in detail.

[0152] Figure 3 It is a schematic flowchart of an audio processing method provided by this application. As Figure 3 shown, this method can include S301 and S302.

[0153] S301, the terminal device obtains the audio data to be processed.

[0154] The architecture of the terminal device can be as Figure 1 and Figure 2 shown.

[0155] As an example, the audio data to be processed can be the audio data obtained by recording.

[0156] Optionally, the audio data to be processed can be obtained by the terminal device itself. For example, the user can input a recording instruction on the terminal device, and the terminal device can record based on the recording instruction. After the recording is completed, the recorded audio data can be obtained. The recorded audio data obtained by the terminal device can include the audio data to be processed.

[0157] Optionally, the audio data to be processed can also be obtained from other devices. For example, the user can input a recording instruction on another device, and the other device can record based on the recording instruction. After the recording is completed, the recorded audio data can be obtained. After the other device obtains the recorded audio data, it can send the recorded audio data to the terminal device. Correspondingly, the terminal device can receive the recorded audio data. The recorded audio data received by the terminal device can include the audio data to be processed.

[0158] S302. The terminal device outputs first information based on the audio data to be processed, and the first information indicates the data density of the audio data to be processed in each of multiple time periods.

[0159] Among them, these multiple time periods can be preset in advance, and the duration of each of these multiple time periods can also be preset in advance. For example, the start time and end time of each time period can be preset in advance.

[0160] Optionally, the duration of each of these multiple time periods can be equal.

[0161] Optionally, the terminal device can first determine these multiple time periods, then parse the audio data to be processed to obtain the data density of the audio data to be processed in each of these multiple time periods, and then output the first information.

[0162] In this method, the data density can represent the amount of data per unit duration. When the duration of each of these multiple time periods is equal, the data density of the audio data to be processed in each time period can be characterized by the amount of audio data of the audio data to be processed in each time period.

[0163] As an example, assume that these multiple time periods can include a first time period and a second time period, the duration of the first time period is equal to the duration of the second time period, the data density of the audio data to be processed in the first time period can be characterized by the amount of data of the audio data to be processed in the first time period, and the data density of the audio data to be processed in the second time period can be characterized by the amount of data of the audio data to be processed in the second time period.

[0164] Assume that the duration of the first time period is equal to the duration of the second time period. If the amount of data of the audio data to be processed in the first time period is greater than the amount of data of the audio data to be processed in the second time period, it means that the data density of the audio data to be processed in the first time period is higher than the data density of the audio data to be processed in the second time period, that is, the audio data of the audio data to be processed in the first time period is denser than the audio data of the audio data to be processed in the second time period.

[0165] In this method, the output of the first information by the terminal device can be understood as: the terminal device displays the first information, or sends the first information to other devices, and after receiving the first information, the other devices display the first information.

[0166] Optionally, the first information (or the data density of each period in these multiple periods) can be displayed in one or more of the following ways: energy bar, color.

[0167] As an example, when the data density is displayed by an energy bar, the more energy bars there are in the period with a higher data density. Or rather, the more energy bars there are in the period with a larger amount of data.

[0168] Suppose the energy bars for different periods can be as Figure 4 shown. Suppose these multiple periods include period 1 and period 2, and the durations of period 1 and period 2 are equal, and the amount of data of the audio data to be processed in period 1 is greater than the amount of data of the audio data to be processed in period 2, then the energy bar of period 1 is more than the energy bar of period 2, or rather, the length of the energy bar of period 1 is greater than the length of the energy bar of period 2. In this example, the shaded part below the progress bar of each period can represent the energy bar of each period. In this example, the data density of each period in these multiple periods is displayed below the audio progress bar.

[0169] Optionally, in this example, the terminal device or other devices can also display the data density of each period in these multiple periods on the audio progress bar.

[0170] Optionally, the energy bars for different periods can also be as Figure 5 shown. Suppose these multiple periods include period 1 and period 2, and the durations of period 1 and period 2 are equal, and the amount of data of the audio data to be processed in period 1 is greater than the amount of data of the audio data to be processed in period 2, then the energy bar of period 1 is more than the energy bar of period 2.

[0171] In this example, the terminal device or other devices can display the data density of each period in these multiple periods below the audio progress bar.

[0172] As another example, when the data density is displayed by color, the colors of the periods with different data densities can be different. For example, the darker the color of the period with a higher data density.

[0173] In this example, the terminal device or other devices can display the data density of each period in these multiple periods on the audio progress bar. The color of each period can be the color corresponding to the position of each period in the progress bar of the audio to be processed.

[0174] Such as Figure 6As shown, assume that these multiple time periods include time period 1 and time period 2, and the durations of time period 1 and time period 2 are equal. If the data volume of the audio data to be processed in time period 1 is greater than the data volume of the audio data to be processed in time period 2, then the color of the position of time period 1 in the progress bar is different from the color of the position of time period 2 in the progress bar.

[0175] Assume that the shaded part in the progress bar can represent red and the blank part can represent yellow. Then, the shaded part can represent the color of time period 1 in the progress bar, and the blank part can represent the color of time period 2 in the progress bar. The color of the position of time period 1 in the progress bar is darker than the color of the position of time period 2 in the progress bar.

[0176] Optionally, when the durations of each of these multiple time periods are equal, the data density of each of these multiple time periods can also be represented by the data volume of each time period.

[0177] As an example, the data volume of each time period can be represented by the bit rate or bit value of the audio data of each time period.

[0178] Among them, the data volume of each time period can be represented in various ways, and this application does not limit this. As an example, the data volume of each time period can be represented by a bar chart, a line chart, or a table, etc.

[0179] Taking the bar chart as an example, the data volume of each time period can be as Figure 7 shown. In this example, assume that these multiple time periods include time period 1, time period 2, time period 3, and time period 4, and the data volume of each of these four time periods is represented by the bit rate of the audio data of each time period.

[0180] In this example, the terminal device or other devices can display the data volume of each of these multiple time periods below the audio progress bar.

[0181] In some possible implementation manners, the terminal device can also obtain second information, and this second information can indicate the audio scene of the audio data to be processed. Then, based on the second information, a first video is generated, and the first video can include a virtual scene corresponding to the audio scene of the audio data to be processed.

[0182] Taking the audio data obtained by recording as the audio data to be processed as an example, the audio scene of this audio data to be processed can be a recording scene, or the real scene during recording.

[0183] In this example, the virtual scene corresponding to the audio scene of the audio data to be processed can be a virtual scene that simulates the real scene during recording. The content included in the virtual scene can be the same as or different from the content in the real scene. For example, as long as the scene types of the audio scene and the virtual scene are the same, such as both the audio scene and the virtual scene are of the type of meeting room, etc.

[0184] In this method, the second information can be input by the user. For example, after the user clicks on the first application on the terminal device, the terminal device can display the first user interface. Then the user can open the audio file corresponding to the audio data to be processed in the first user interface and input the second information. The audio file contains the audio data to be processed.

[0185] Among them, the first application can be an application for generating a video based on the audio data. As an example, when the audio data is recorded audio data, the first application can be a recording application or other applications.

[0186] As an example, the first user interface can be as Figure 8 shown. The first user interface includes a button for selecting an audio file and a button for inputting an audio scene.

[0187] After the user clicks the button for selecting an audio file in the first user interface, the terminal device can display a second user interface as Figure 9 shown. At least one audio file can be displayed in the second user interface. In this example, the at least one audio file can include Audio File 1, Audio File 2, and Audio File 3.

[0188] Suppose the user clicks on Audio File 1 in the second user interface, the terminal device can return to the first user interface. At this time, the first user interface can be as Figure 10 shown, and Audio File 1 can be displayed in the first user interface.

[0189] In this example, after the user selects the audio file, the user can click the button for inputting an audio scene in the first user interface as Figure 10 shown. After the user clicks the button for inputting an audio scene, the terminal device can obtain the second information.

[0190] Optionally, the user can also perform operations such as deleting the selected audio file and selecting an audio file in the first user interface as Figure 10 shown.

[0191] It can be understood that Figures 8 to 10 the user interface shown is only an example and does not limit the scope of this application. As long as the terminal device can obtain the second information.

[0192] In a possible implementation, after the terminal device obtains the second information, it can generate a virtual scene corresponding to the audio scene through a scene generation model.

[0193] Among them, the scene generation model can be an existing model.

[0194] As an example, assume that when the user inputs "conference room" in the first user interface as shown in Figure 10 , the terminal device can generate a virtual conference room scene based on the scene generation model, and the virtual conference room scene can include virtual items such as tables, chairs, projectors, etc.

[0195] Optionally, the terminal device can also perform semantic analysis on the audio data to be processed, determine some items that may exist in the audio scene, such as notebooks, cups, etc., and add virtual items corresponding to these items in the virtual conference scene.

[0196] In another possible implementation, the terminal device can also preset many audio scenes in advance, and select a corresponding audio scene from the preset audio scenes based on the second information input by the user to generate a virtual scene.

[0197] As an example, assume that when the user inputs "conference room" in the first user interface as shown in Figure 10 , the terminal device can select the conference room scene from the preset audio scenes to generate a virtual scene.

[0198] Optionally, the audio data to be processed can include audio data of one or more speakers. The virtual scene corresponding to the audio scene of the audio data to be processed can include virtual characters corresponding to at least one of the one or more speakers.

[0199] Among them, the virtual characters corresponding to each of the at least one speaker can be virtual characters that simulate each speaker.

[0200] In this method, the terminal device can determine that the audio data to be processed includes audio data of one or more speakers through voice recognition technology, then determine the position of at least one of the one or more speakers in the audio scene through sound source localization technology, and then determine the position of the virtual character corresponding to each of the at least one speaker in the virtual scene based on the position of each of the at least one speaker in the audio scene.

[0201] Among them, the voice recognition technology here can be an existing technology and will not be elaborated here.

[0202] Optionally, the terminal device can also record the speaking period of each of the one or more speakers.

[0203] In a possible implementation, the position of each speaker among the one or more speakers in the audio scene can be determined during the process of generating the audio.

[0204] Taking the audio data to be processed as the audio data obtained by recording as an example, during the recording process, the recording device can locate each speaker among the one or more speakers based on the sound source localization technology to determine the position of each speaker in the audio scene.

[0205] Optionally, the recording device can be the terminal device or other devices other than the terminal device.

[0206] When the recording device is the terminal device, during the recording process, the terminal device can determine the position of each speaker among the one or more speakers in the audio scene based on the sound source localization technology.

[0207] When the recording device is other devices, during the recording process, the other devices can determine the position of each speaker among the one or more speakers in the audio scene based on the sound source localization technology. When the other devices send the audio data to the terminal device, they can also send the position information to the terminal device, and the position information includes the position of each speaker among the one or more speakers in the audio scene.

[0208] Optionally, the terminal device can also record the position of each speaker among the one or more speakers at each moment among multiple moments.

[0209] Alternatively, the terminal device can also record the position of each speaker among the one or more speakers in each period among multiple periods. Wherein, the position of each speaker in each period remains unchanged, and the position of at least one speaker in different periods is different or has changed.

[0210] In this method, after the terminal device obtains the position of each speaker among the one or more speakers in the audio scene, it can load the virtual characters corresponding to each speaker in the virtual scene based on the position of each speaker among the one or more speakers in the audio scene.

[0211] As an example, the terminal device can directly load the virtual characters corresponding to each speaker among the one or more speakers into the virtual scene.

[0212] In this example, each video frame in the first video can include the virtual characters corresponding to each speaker among the one or more speakers.

[0213] As another example, the terminal device can load the virtual characters corresponding to each speaker into the virtual scene based on the speaking period of each speaker among the one or more speakers.

[0214] For example, assume that the one or more speakers include Speaker 1 and Speaker 2. Among them, Speaker 1 speaks in Period 1 and Speaker 2 speaks in Period 2. Then, each video frame in the first video located in the first period includes the virtual character corresponding to Speaker 1, and each video frame in the first video located in the second period includes the virtual character corresponding to Speaker 2.

[0215] In this example, each video frame in the first video may include the virtual characters corresponding to at least one of the one or more speakers.

[0216] In this method, the position of the virtual character corresponding to each speaker among at least one speaker in the virtual scene may be consistent with the position of each speaker in the audio scene.

[0217] For example, the relative position of the virtual character corresponding to each speaker in the virtual scene may be consistent with the relative position of each speaker in the audio scene, or the relative position of the virtual characters corresponding to any two speakers in the virtual scene may be consistent with the relative position of these two speakers in the audio scene.

[0218] Another example is that the position of the virtual character corresponding to each speaker in the virtual scene changes as the position of each speaker in the audio scene changes.

[0219] Assume that when the terminal device determines that the position of any one speaker changes, it may update the position of the virtual character corresponding to the any one speaker in the virtual scene based on the position information of the any one speaker.

[0220] It can be understood that the consistency here is not completely consistent in the strict sense. As long as the position of the virtual character corresponding to each speaker among at least one speaker in the virtual scene is basically consistent with the position of each speaker in the audio scene, or the position difference between the position of the virtual character corresponding to each speaker among at least one speaker in the virtual scene and the position of each speaker in the audio scene is less than or equal to the preset distance threshold.

[0221] In this method, the position of the virtual character corresponding to each speaker among at least one speaker in the virtual scene is consistent with the position of each speaker in the audio scene, making the generated virtual scene closer to the audio scene and conducive to improving the user experience.

[0222] Optionally, after the terminal device generates the first video, it may also output the first information.

[0223] In this method, the output of the first information by the terminal device can be understood as: after generating the first video, the terminal device displays the first information, or sends the first video and the first information to other devices, and after receiving the first video and the first information, the other devices display the first video and the first information.

[0224] As an example, the first information can be displayed in the first video or below the first video.

[0225] For example, assuming the display mode of the first information is as Figure 4 shown, the terminal device or other devices can display the first information on or below the progress bar of the first video.

[0226] Another example, assuming the display mode of the first information is as Figure 5 shown, the terminal device or other devices can display the first information below the progress bar of the first video.

[0227] Another example, assuming the display mode of the first information is as Figure 6 shown, the terminal device or other devices can display the first information below the progress bar of the first video.

[0228] Another example, assuming the display mode of the first information is as Figure 7 shown, the terminal device or other devices can display the first information below the progress bar of the first video.

[0229] In some possible implementation manners, the terminal device can also play the first video, and the picture focus of the first video in the first period among multiple periods is on the virtual character corresponding to the first speaker among one or more speakers, and the first period can be the period when the first speaker is speaking.

[0230] Wherein, the first speaker can include one or more speakers.

[0231] The picture focus of the first video in the first period being on the virtual character corresponding to the first speaker can be understood as: in each video frame of the first video located in the first period, the proportion of the picture of the virtual character of the first speaker in each video frame is greater than the proportion of the picture of the virtual characters of other speakers except the first speaker in each video frame.

[0232] In this method, when generating the first video, the terminal device can place the picture focus of each period among multiple periods on the virtual character corresponding to the speaker in each period.

[0233] As an example, assume that the speaker is the same or remains unchanged during the same period, and different speakers are different or have changed during different periods. The terminal device can first record the positions of the speakers in the audio scene in each period among multiple periods, and then determine the positions of the virtual characters corresponding to the speakers in each period in the virtual scene based on the positions of the speakers in the audio scene in each period. Finally, place the focus of the picture in each period on the positions of the virtual characters corresponding to the speakers in each period in the virtual scene.

[0234] Optionally, the position of the virtual character in the virtual scene can be represented by a recognition frame. The terminal device can determine the position of the virtual character in the virtual scene based on the position of the recognition frame in the virtual scene.

[0235] Correspondingly, the terminal device placing the focus of the picture on the position of any virtual character in the virtual scene can be understood as: placing the focus of the picture on the position of the center point of the recognition frame of the any virtual character in the virtual scene.

[0236] For example, assume that the first period is the speaking period of the first speaker. Then, when the terminal device generates the first video, it can place the focus of the picture in the first period on the virtual character corresponding to the first speaker.

[0237] Specifically, the terminal device can first determine that the first period is the speaking period of the first speaker, and then determine the position of the recognition frame of the virtual character corresponding to the first speaker in the virtual scene based on the position of the first speaker in the audio scene in the first period. Then, in the first period, place the focus of the picture on the position of the center point of the recognition frame in the virtual scene.

[0238] Taking the first speaker including one speaker as an example, assume that the first speaker starts speaking at the first moment. Then, before the first moment, the pictures of the first speaker at different moments in the video can be as Figure 11 shown in (a). And assume that the first speaker ends speaking at the second moment. Then, during the period between the first moment and the second moment, the picture of the first speaker in the video can be as Figure 11 shown in (b). After the second moment, the picture of the first speaker in the video can be as Figure 11 shown in (c). In this example, the period between the first moment and the second moment can be the first period.

[0239] In this method, the terminal device can enlarge the proportion of the picture of the speaker in the first video, which can provide a more realistic audio-visual experience for the user and is beneficial to improving the user experience.

[0240] Optionally, the first video may also display the action information of the virtual character corresponding to the first speaker during the first time period. The action information of the virtual character corresponding to the first speaker may include the mouth actions of the virtual character corresponding to the first speaker.

[0241] In this method, when generating the first video, the terminal device may also add action information to the virtual characters corresponding to the speakers in each time period among multiple time periods.

[0242] As an example, assuming that the speakers in the same time period are the same or remain unchanged, and the speakers in different time periods are different or have changed, the terminal device may first record the positions of the speakers in each time period in the audio scene, then determine the positions of the virtual characters corresponding to the speakers in each time period in the virtual scene based on the positions of the speakers in the audio scene in each time period, and finally add the action information of the virtual characters at the positions of the virtual characters corresponding to the speakers in each time period in the virtual scene.

[0243] Optionally, the position of the virtual character in the virtual scene may be represented by a recognition frame. The terminal device may determine the position of the virtual character in the virtual scene based on the position of the recognition frame in the virtual scene.

[0244] Correspondingly, the terminal device adding the action information of the virtual character at the position of the virtual character in the virtual scene may be understood as: adding the action information of the virtual character in the recognition frame of the virtual character.

[0245] For example, assuming that the first time period is the speaking time period of the first speaker, then when generating the first video, the terminal device may add action information to the virtual character corresponding to the first speaker during the first time period.

[0246] Specifically, the terminal device may first determine that the first time period is the speaking time period of the first speaker, then determine the position of the recognition frame of the virtual character corresponding to the first speaker in the virtual scene based on the position of the first speaker in the audio scene during the first time period, and then add the action information of the virtual character in the recognition frame of the virtual character corresponding to the first speaker during the first time period.

[0247] In this method, the first video may include the action information of the virtual characters corresponding to the speakers, which can provide a more realistic audio-visual experience for the user and is beneficial to improving the user experience.

[0248] In some possible implementation manners, the terminal device may also obtain third information, and the third information may instruct the user to click on the virtual character corresponding to the second speaker among one or more speakers, and then amplify the voice of the second speaker.

[0249] As an example, the terminal device can record the positions of the virtual characters corresponding to the speakers in each of multiple time periods in a virtual scene. After the terminal device obtains the third information, it can first determine the virtual character clicked by the user based on the time period when the user obtains the third information and the position clicked by the user, then determine the speaker corresponding to the virtual character, and finally select the voice of the speaker from the audio data to be processed and amplify the voice of the speaker.

[0250] Optionally, the terminal device can also reduce the voices of other speakers except the second speaker.

[0251] As an example, the terminal device can record the positions of the virtual characters corresponding to the speakers in each of multiple time periods in a virtual scene. After the terminal device obtains the third information, it can first determine the virtual character clicked by the user based on the time period when the user obtains the third information and the position clicked by the user, then determine the speaker corresponding to the virtual character, and finally select the voices of other speakers except the speaker from the audio data to be processed and reduce the voices of the other speakers.

[0252] Among them, the second speaker can be any one or more of the one or more speakers.

[0253] Taking the second speaker including one speaker as an example, assuming that a certain video frame includes virtual characters corresponding to two speakers, after the user clicks on the virtual character corresponding to one of the speakers, the terminal device can amplify the voice of the speaker and can reduce the voice of the other speaker.

[0254] In this method, the terminal device can amplify the voice of the second speaker based on the third information, which is beneficial for the user to accurately obtain the target information required by the user from the amplified voice.

[0255] Optionally, assuming that the user clicks on the virtual character corresponding to the second speaker at the third moment, and the user clicks on the virtual character corresponding to other speakers except the second speaker at the fourth moment, the time period between the third moment and the fourth moment can be the second time period, then the terminal device can place the picture focus of the first video on the virtual character corresponding to the second speaker during the second time period.

[0256] Optionally, the method for the terminal device to place the picture focus of the second time period on the virtual character corresponding to the second speaker can refer to the method for placing the picture focus of the first time period on the virtual character corresponding to the first speaker in the foregoing embodiments, which will not be elaborated here.

[0257] Such as Figure 12As shown in (a) in [reference], assume that a video frame includes virtual characters corresponding to two speakers. After the user clicks on the virtual character corresponding to one of the speakers, the video frame can be as shown in Figure 12 (b) in [reference].

[0258] In this method, the terminal device can also increase the proportion of the screen occupied by the second speaker in the first video, which can provide a more realistic audio-visual experience for the user and is beneficial to improving the user experience.

[0259] In some possible implementation manners, the terminal device can also determine the time periods in which the audio data volume exceeds a preset data volume threshold from multiple time periods, and then mark the time periods in which the audio data volume exceeds the preset data volume threshold.

[0260] As an example, the time periods in which the audio data volume exceeds the preset data volume threshold can be marked by highlighting. Optionally, it can be marked in the audio progress bar, or the first information, or the progress bar of the first video.

[0261] Taking the example of marking the time periods in which the audio data volume exceeds the preset data volume threshold in the progress bar of the first video, assume that the terminal device can record the positions of each time period in the progress bar of the first video in multiple time periods. After the terminal device determines the time periods in which the audio data volume exceeds the preset data volume threshold from multiple time periods, it can first determine the position of the start moment of the time period in which the audio data volume exceeds the preset data volume threshold in the progress bar of the first video, and then mark at the position of the time period in which the audio data volume exceeds the preset data volume threshold in the progress bar of the first video.

[0262] Assume that the first video can be as shown in Figure 13 wherein, the shaded part in the progress bar of the first video can represent highlighting, and the blank part in the progress bar of the first video can represent non-highlighting. The shaded part represents the time periods in which the audio data volume exceeds the preset data volume threshold, and the blank part represents the time periods in which the audio data volume does not exceed the preset data volume threshold.

[0263] In this method, by marking the time periods in which the audio data volume exceeds the preset data volume threshold, the user can determine the time periods in which the audio data volume exceeds the preset data volume threshold in the audio or video, and then obtain the target information from the time periods in which the audio data volume exceeds the preset data volume threshold, which is beneficial for the user to quickly obtain the target information.

[0264] Furthermore, the terminal device can also perform semantic analysis on the audio data within the time periods in which the audio data volume exceeds the preset data volume threshold, obtain the key information within the time periods in which the audio data volume exceeds the preset data volume threshold, and output the key information.

[0265] Optionally, the key information may be displayed at the start time of the period when the audio data volume exceeds a preset data volume threshold. As an example, the key information may be a keyword.

[0266] Optionally, the key information may be displayed in the audio progress bar, or the first information, or the progress bar of the first video.

[0267] Taking the display of the key information in the progress bar of the first video as an example, assuming that the terminal device can record the position of each period in the progress bar of the first video in multiple periods, after the terminal device determines the period when the audio data volume exceeds the preset data volume threshold from multiple periods, it can first determine the position of the start time of the period when the audio data volume exceeds the preset data volume threshold in the progress bar of the first video, and then display the key information at the position of the start time of the period when the audio data volume exceeds the preset data volume threshold in the progress bar of the first video.

[0268] Assume that the first video can be as Figure 14 shown, where the shaded part in the progress bar of the first video may represent the period when the audio data volume exceeds the preset data volume threshold, and the key information is displayed at the start time of each period corresponding to the shaded part.

[0269] Optionally, in some embodiments, it may not be necessary to display the period when the audio data volume exceeds the preset data volume threshold, and it is only necessary to display the key information at the start time of the period when the audio data volume exceeds the preset data volume threshold.

[0270] In this method, by displaying the key information, the user can determine the period containing the target information required by the user from at least one period when the audio data volume exceeds the preset data volume threshold, and then obtain the target information from the audio or video in the period containing the target information required by the user, which is beneficial for the user to quickly obtain the target information.

[0271] Figure 15 Schematic diagram of an audio processing device provided by an embodiment of the present application. As Figure 15 shown, the audio processing device 1500 may include an acquisition module 1501 and a processing module 1502.

[0272] As an example, the audio processing device 1500 may be used to implement the Figure 3 method of the embodiment shown. Among them, the acquisition module 1501 may be used to execute S301, and the processing module 1502 may be used to execute S302.

[0273] Figure 16 Schematic diagram of an audio processing device provided by another embodiment of the present application. As Figure 16As shown, the audio processing device 1600 includes a processor 1601 and an interface circuit 1602. The processor 1601 and the interface circuit 1602 are coupled to each other. It can be understood that the interface circuit 1602 can be a transceiver or an input / output interface. Optionally, the audio processing device 1600 may further include a memory 1603 for storing instructions executed by the processor 1601 or input data required for the processor 1601 to execute instructions or data generated after the processor 1601 executes instructions.

[0274] As an example, the processor 1601 can be used to implement the functions of the above-mentioned processing module 1502, and the interface circuit 1602 can be used to implement the functions of the above-mentioned acquisition module 1501.

[0275] In this example, the audio processing device 1600 can be a terminal device or a chip applied to a terminal device.

[0276] The method steps in the embodiments of the present application can be implemented in a hardware manner or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, a register, a hard disk, a removable hard disk, a CD-ROM, or any other form of storage medium well-known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC. Additionally, the ASIC can be located in a network device or a terminal device. Of course, the processor and the storage medium can also exist as discrete components in a network device or a terminal device.

[0277] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are executed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device, or other programmable devices. The computer program or instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program or instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; it can also be an optical medium, such as a digital video disc; or it can be a semiconductor medium, such as a solid-state drive.

[0278] In various embodiments of the present application, if there is no special description and logical conflict, the terms and / or descriptions between different embodiments are consistent and can be referenced to each other. The technical features in different embodiments can be combined to form new embodiments according to their internal logical relationships.

[0279] It can be understood that the various numerical numbers involved in the embodiments of the present application are only for the convenience of description and do not limit the scope of the embodiments of the present application. The magnitudes of the serial numbers of the above processes do not mean the order of execution, and the order of execution of each process should be determined by its function and internal logic.

Claims

1. An audio processing method, characterized in that: The method comprises: Get the audio data to be processed; First information is output based on the audio data to be processed, where the first information indicates a data density of each time period of the audio data to be processed in a plurality of time periods.

2. The method according to claim 1, characterized in that The method further comprises: Acquire second information, where the second information indicates an audio scene of the audio data to be processed; A first video is generated based on the second information, wherein the first video includes a virtual scene corresponding to the audio scene of the audio data to be processed.

3. The method according to claim 2, characterized in that The audio data to be processed includes audio data of one or more speakers, and the virtual scene includes a virtual character corresponding to at least one of the one or more speakers.

4. The method according to claim 3, characterized in that The position of the virtual character corresponding to each speaker of the at least one speaker in the virtual scene is consistent with the position of each speaker in the audio scene.

5. The method according to any one of claims 2 to 4, characterized in that The method further comprises: The first video is played, wherein a screen focus of the first video in a first time period among the multiple time periods is on a virtual character corresponding to a first speaker among one or more speakers, and the first time period is a time period in which the first speaker speaks.

6. The method according to claim 5, characterized in that The first video displays action information of the virtual character corresponding to the first speaker in the first time period, and the action information of the virtual character corresponding to the first speaker includes mouth movements of the virtual character corresponding to the first speaker.

7. The method according to claim 5 or 6, characterized in that: The method further comprises: Acquire third information, where the third information indicates that the user clicks on the virtual character corresponding to the second speaker among the one or more speakers; Amplify the second speaker's voice.

8. The method according to any one of claims 5 to 7, characterized in that The screen focus of the first video in the second time period of the multiple time periods is on the virtual character corresponding to the second speaker among the one or more speakers, the starting time of the second time period is the time when the user clicks on the virtual character corresponding to the second speaker, and the ending time of the second time period is the time when the user clicks on other speakers among the one or more speakers except the second speaker.

9. The method according to any one of claims 1 to 8, characterized in that The data density of each time period in the multiple time periods is displayed in one or more of the following ways: energy bar, color.

10. The method according to any one of claims 1 to 9, characterized in that The method further comprises: Determining a time period in which the amount of audio data exceeds a preset data amount threshold from the multiple time periods; The time period in which the audio data volume exceeds a preset data volume threshold is marked.

11. The method according to claim 10, characterized in that The method further comprises: Performing semantic analysis on the audio data in the time period when the amount of audio data exceeds the preset data amount threshold, to obtain key information in the time period when the amount of audio data exceeds the preset data amount threshold; The key information is output.

12. An audio processing device, characterized in that: The method comprises a functional module for implementing the method according to any one of claims 1 to 11.

13. An audio processing device, characterized in that: include: Memory and processor; The memory is used to store program instructions; The processor is configured to execute program instructions in the memory to implement the method according to any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program code for computer execution, wherein the program code includes instructions for implementing the method according to any one of claims 1 to 11.

15. A computer program product, characterized in that The computer program product comprises instructions for implementing the method according to any one of claims 1 to 11.