Audio data processing method and device, equipment and readable medium
By analyzing and discarding audio frames without sound information in real time during the voice conversation, and generating a recording file, the problem of low proportion of effective information in the recording file is solved, and the recording effect and efficiency are improved.
Patent Information
- Application Number
- CN202410168923.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-06
- Publication Date
- 2025-08-08
AI Technical Summary
During audio recording, long-term no one speaks or intermittent sounds lead to a small proportion of effective information in the recording file. The existing manual start and stop recording method is inconvenient and easy to operate incorrectly, affecting the recording effect.
During the voice session, the audio data is obtained in real time and audio frame analysis is performed, the soundless information frame is discarded, and the sound information frame is retained to generate the recording file.
It increases the proportion of effective information in the recorded file, improves the recorder's user experience and processing efficiency, and reduces the resource occupancy of invalid information.
Smart Images

Figure CN120452483A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computers and communication technologies, and in particular to an audio data processing method, a data push device, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] With the development of the internet, audio interaction has become a ubiquitous part of everyday life, with applications such as voice conferencing and live streaming. These scenarios often require audio recording. However, if there are long periods of silence or intermittent audio during recording, the resulting audio will contain very little information that is truly useful to the recorder, hindering subsequent use. Therefore, improving the proportion of useful information in audio recordings is a pressing issue. Summary of the Invention
[0003] The embodiments of the present application provide an audio data processing method, an audio data processing device, an electronic device, a computer-readable storage medium, and a computer program product, which can increase the proportion of effective information in an audio recording file.
[0004] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by practice of the present application.
[0005] According to one aspect of an embodiment of the present application, a method for processing audio data is provided, the method comprising:
[0006] During the voice conversation, obtaining the generated conversation audio data;
[0007] For each audio frame in the conversation audio data, performing audio analysis on the audio frame to obtain an analysis result;
[0008] If the analysis result indicates that the audio frame does not contain sound information, the audio frame is discarded, and an audio recording file corresponding to the voice conversation is generated based on the remaining audio frames in the conversation audio data.
[0009] According to one aspect of an embodiment of the present application, there is provided an audio data processing device, the device comprising an acquisition unit and a processing unit, wherein:
[0010] The acquisition unit is used to acquire the conversation audio data generated during the voice conversation;
[0011] The processing unit is configured to perform audio analysis on each audio frame in the conversation audio data to obtain an analysis result;
[0012] The processing unit is further configured to discard the audio frame if the analysis result indicates that the audio frame does not contain sound information, and generate an audio recording file corresponding to the voice conversation based on the remaining audio frames in the conversation audio data.
[0013] According to one aspect of an embodiment of the present application, an embodiment of the present application provides an electronic device, comprising one or more processors; a storage device for storing one or more computer programs, wherein when the one or more computer programs are executed by the one or more processors, the electronic device implements the audio data processing method as described above.
[0014] According to one aspect of an embodiment of the present application, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor of an electronic device, the electronic device executes the audio data processing method as described above.
[0015] According to one aspect of an embodiment of the present application, an embodiment of the present application provides a computer program product, including a computer program, wherein the computer program is stored in a computer-readable storage medium, and a processor of an electronic device reads and executes the computer program from the computer-readable storage medium, so that the electronic device performs the audio data processing method as described above.
[0016] In the technical solution provided in the embodiment of the present application, by performing audio analysis on the audio frames contained in the conversation audio data, it is possible to accurately identify which audio frames do not contain sound information and are therefore discarded, and which audio frames contain sound information and are therefore recorded; the audio recording files obtained in this way all contain sound information, and sound information can convey effective information in the voice conversation, thus greatly increasing the proportion of effective information in the audio recording files. In addition, the embodiment of the present application obtains the conversation audio data generated during the voice conversation and performs audio analysis synchronously. This method allows the recorder to obtain an audio recording file with a high proportion of effective information in a relatively short period of time after the voice conversation ends, which is beneficial to improving the recorder's user experience. At the same time, it also avoids audio frames containing invalid information from occupying more resources, which is beneficial to improving the processing efficiency of audio data.
[0017] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings are incorporated into and constitute a part of the specification, illustrating embodiments consistent with the present application and, together with the specification, serving to explain the principles of the present application. Obviously, the drawings described below are merely some embodiments of the present application, and a person of ordinary skill in the art can derive other drawings based on these drawings without inventive effort. In the drawings:
[0019] Figure 1 This is a structural diagram of an audio data processing system provided by an embodiment of the present application;
[0020] Figure 2 This is a flowchart of an audio data processing method provided by an embodiment of the present application;
[0021] Figure 3 This is a schematic diagram of an audio recording process provided by an embodiment of the present application;
[0022] Figure 4 This is a flowchart of another audio data processing method provided by an embodiment of the present application;
[0023] Figure 5 This is a schematic diagram of the data volume of a data slice provided in an embodiment of the present application;
[0024] Figure 6 This is a schematic diagram of a VAD detection process provided by an embodiment of the present application;
[0025] Figure 7 This is a schematic diagram of a music detection process provided by an embodiment of the present application;
[0026] Figure 8 This is a schematic diagram of an audio analysis process provided by an embodiment of the present application;
[0027] Figure 9 is a structural block diagram of an audio data processing device shown in an exemplary embodiment of the present application;
[0028] Figure 10 A schematic diagram of the structure of a computer system suitable for implementing an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0029] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numerals in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0030] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically separate entities. That is, these functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0031] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations, nor must they be executed in the order described. For example, some operations may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.
[0032] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0033] It should also be noted that the term "plurality" used in this application refers to two or more. "And / or" describes the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.
[0034] With the development of the internet, audio interaction has become ubiquitous in everyday life, such as voice conferencing and interactive live streaming. Audio interaction scenarios often require recording. If no one speaks for extended periods of time during recording, or if the audio is intermittent, the resulting recorded audio will contain very little information that is truly useful to the recorder, hindering subsequent use.
[0035] For these types of audio interaction scenarios, the current approach is to increase the proportion of effective information in the final recorded audio by manually turning off the recording when no sound is heard and manually starting it again when sound is heard. However, this manual turning-off and starting method forces the recorder to manually start and stop the recording during the audio interaction, which is inconvenient and prone to misoperation, resulting in poor audio quality.
[0036] Based on this, an embodiment of the present application provides an audio data processing solution that obtains the conversation audio data generated by the voice conversation during the voice conversation and performs audio analysis on each audio frame in the obtained conversation audio data to obtain the corresponding analysis results. If the analysis result of an audio frame indicates that the audio frame does not contain sound information, the audio frame is discarded; if the analysis result of an audio frame indicates that the audio frame contains sound information, an audio recording file corresponding to the voice conversation is generated based on the audio frame.
[0037] Among them, voice conversations can specifically include voice calls, interactive live broadcasts, online meetings, and other audio interaction scenarios in which one or more conversation partners join and sound needs to be emitted.
[0038] The acquired conversation audio data may be conversation audio data generated at the current moment during the voice conversation, or historical conversation audio data generated within a period of time before the current moment, which is not limited here.
[0039] Furthermore, the audio analysis performed on the audio frames contained in the acquired conversation audio data primarily analyzes whether the audio frames contain sound information. Sound information can specifically refer to sounds produced by any person or object; alternatively, it can refer to sounds produced by a specified object. For example, if the specified object is a person, the sound information specifically refers to a human voice.
[0040] It is not difficult to see that this solution can accurately identify which audio frames do not contain sound information and are discarded, and which audio frames contain sound information and are recorded by performing audio analysis on the audio frames contained in the conversation audio data; the audio recording files recorded in this way all contain sound information, and sound information can convey effective information in the voice conversation, so this solution is conducive to increasing the proportion of effective information in the audio recording files.
[0041] In addition, this solution obtains the conversation audio data generated during the voice conversation and performs audio analysis synchronously. Compared with the method of first recording all the conversation audio data generated during the voice conversation, then performing audio analysis and editing the audio based on the analysis results, this solution can reuse the time during the voice conversation, allowing the recorder to obtain an audio recording file with a high proportion of valid information in a short period of time after the voice conversation ends, which is beneficial to improving the recorder's user experience. At the same time, this solution will promptly discard audio frames that do not contain valid information during the voice conversation, avoiding audio frames with invalid information from occupying a large amount of resource space, which is beneficial to improving the processing efficiency of audio data.
[0042] Based on the above audio data processing solution, the present application embodiment provides an audio data processing system, which can be found in Figure 1 , Figure 1 The audio data processing system shown may include a terminal device 110 and a server 120. The number of terminal devices 110 may include multiple, such as terminal device 111, terminal device 112, etc., and the number of servers 120 may include multiple. A communication connection is established between any terminal device and any server. For example, the terminal device 110 may include any one or more of a smartphone, a tablet computer, a laptop computer, a desktop computer, an intelligent voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, and a smart wearable device. Various types of clients such as a live broadcast client, a social client, an online conference client, a multimedia playback client, a map client, etc. may be installed in the terminal device 110. The server 120 may be a server or server cluster that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Any terminal device 110 and any server 120 may be directly or indirectly connected in communication via wired or wireless communication, and this application does not limit this.
[0043] In some embodiments, the above audio data processing method can be performed by Figure 1 The terminal device 110 in the audio data processing system shown is executed, and the specific execution process is as follows: the terminal device 110 can obtain the conversation audio data generated by the voice conversation during the voice conversation. Then, the terminal device 110 can perform audio analysis on each audio frame in the obtained conversation audio data to obtain the analysis result of the audio frame. Thereafter, if it is detected that the analysis result of the audio frame indicates that the audio frame does not contain sound information, the terminal device 110 can discard the audio frame; if the analysis result of the audio frame indicates that the audio frame does not contain sound information, the terminal device 110 can generate an audio recording file corresponding to the voice conversation based on the audio frame. Furthermore, the terminal device 110 can also output the audio recording file.
[0044] Optionally, the above audio data processing method may also be performed by Figure 1 The server 120 in the audio data processing system shown is executed. The specific execution process can refer to the specific execution process of the terminal device 110, which will not be repeated here.
[0045] In other embodiments, the above audio data interaction method can be run in an audio data interaction system, which can include a terminal device and a server. Figure 1 The audio data interaction system shown is jointly completed by the terminal device 110 and the server 120, and its specific execution process is: each terminal device 110 collects the audio data generated by the object in the voice conversation; each terminal device 110 sends the collected audio data to the server 120; then, the server 120 obtains the conversation audio data of the voice conversation based on the received audio data.
[0046] Afterwards, server 120 may perform audio analysis on each audio frame in the acquired conversation audio data to obtain an analysis result for that audio frame. If the analysis result indicates that the audio frame contains no sound information, server 120 may discard the audio frame. If the analysis result indicates that the audio frame contains no sound information, server 120 may generate an audio recording file corresponding to the voice conversation based on the audio frame. Finally, after the conversation ends, server 120 may send the resulting audio recording file to each terminal device 110.
[0047] It should be noted that the embodiments of the present application can be applied to various scenarios that require audio recording, including but not limited to cloud technology, AI (Artificial Intelligence), smart home, smart transportation, assisted driving and other scenarios, and can also be used for live broadcast applications, online conference applications, social applications and any other applications that require audio recording, without limitation.
[0048] It should be noted that in the specific implementation of this application, if the conversation audio data and other related data or information are related to the object, when the embodiment of this application is applied to a specific product or technology, it is necessary to obtain the object's permission or consent, and the collection, use and processing of the relevant data or information must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0049] The following describes in detail the various implementation details of the technical solutions of the embodiments of the present application:
[0050] like Figure 2 As shown, Figure 2 This is a flow chart of an audio data processing method according to an embodiment of the present invention. The method can be applied to Figure 1The audio data processing system shown in FIG. 1 may be executed by a terminal device or a server, or may be executed by both the terminal device and the server. In the embodiment of the present application, the method is described by taking the method executed by the terminal device as an example. The audio data processing method may include S201 to S203, which are described in detail as follows:
[0051] S201. During a voice conversation, obtain generated conversation audio data.
[0052] In an embodiment of the present application, during a voice conversation, audio data can be collected by a receiver in a terminal device. Since voice conversations continuously generate audio data, the acquired audio data can be either the current audio data or historical audio data generated within a period of time prior to the current time, without limitation. Therefore, the audio data can include one or more audio frames.
[0053] In one embodiment, as mentioned above, a voice conversation may contain multiple conversation objects, and each conversation object in the voice conversation corresponds to a client related to the voice conversation. The client corresponding to the conversation object with the conversation recording function enabled can be referred to as a local client, which can be run on the terminal device that serves as the execution subject in the embodiment of the present application; and the clients corresponding to other conversation objects that join the voice conversation can be referred to as remote clients, which can be run on other terminal devices in the audio data processing system.
[0054] Therefore, the process of obtaining session audio data may include: first, during the voice session, collecting the local audio data generated by the local client in the voice session; then, obtaining the remote audio data generated by the remote client joining the voice session in the voice session; finally, mixing the local audio data and the remote audio data to obtain the session audio data of the voice session.
[0055] The audio mixing process specifically involves merging the audio frames corresponding to the same moment in the local audio data and the remote audio data.
[0056] Optionally, since the audio frames in the local audio data and the audio frames in the remote audio data are arranged according to the time when the audio occurs; therefore, the specific process of the mixing processing may include: for each audio frame in the local audio data, obtaining a data frame in the remote audio data that is arranged in the same order as the audio frame; merging the data frame with the obtained data frame to obtain the conversation audio data.
[0057] The session time of the voice conversation corresponding to the local audio data used in the merging process must be the same as the session time of the voice conversation corresponding to the remote audio data. If the local audio data is audio data collected by the local client within the third to fourth minute of the voice conversation, the remote audio data must be audio data collected by the remote client within the third to fourth minute of the voice conversation.
[0058] Optionally, the local client is also a remote client relative to the remote client. Therefore, each client needs to send the local audio data it collects to other clients.
[0059] Specifically, the local client can perform audio quality enhancement processing on the audio frames contained in the local audio data according to the arrangement order of the audio frames in the local audio data to obtain enhanced audio frames; then, the enhanced audio frames are encoded; whenever the local client detects the encoded enhanced audio frame, the encoded enhanced audio frame is sent to the server, so that every time the server receives the encoded enhanced audio frame, it sends the encoded enhanced audio frame to the remote client.
[0060] Among them, the audio quality enhancement processing may specifically include one or more processing methods for improving voice quality, such as acoustic echo cancellation (AEC), automatic gain control (AGC), and automatic noise suppression (ANS), which are not limited here.
[0061] After receiving the audio frames, the remote client typically plays them directly. However, before sending them, the local client performs audio quality enhancement processing on the audio frames, reducing noise and increasing clarity. This makes the sound played by the remote client clearer, improving the conversational experience for the other party. Similarly, the remote client can also perform audio quality enhancement on the collected remote audio data before uploading it to the server, but this is not discussed here.
[0062] In addition, the streaming audio frame sending method, in which the client sends each encoded enhanced audio frame to the server, and the server sends each encoded enhanced audio frame to the remote client, is conducive to the real-time transmission of audio frames, thereby improving the real-time performance of voice conversations and further improving the conversation experience of the conversation partners.
[0063] S202: Perform audio analysis on each audio frame in the conversation audio data to obtain an analysis result.
[0064] In the embodiment of the present application, the analysis result is used to indicate whether the audio frame contains sound information. The sound information can be any sound produced by any person or object, such as human voice, music, or bird song.
[0065] Optionally, the sound information may also be a sound produced by a specified object. For example, the specified object may be set to a person, and the sound information specifically refers to the sound produced by the person. Optionally, the sound information may also refer to a sound containing specified content. For example, the specified content may be set to 666, and the sound information specifically may be a sound that is the same as or similar to 666. Optionally, the sound information may also include music, etc., which is not limited here.
[0066] In one embodiment, when the sound information is the sound produced by a specified object, the specific process of audio analysis may include: obtaining the sound frequency range corresponding to the specified object and the sound intensity range corresponding to the specified object; if the frequency of the sound contained in the audio frame is within the sound frequency range corresponding to the specified object, and the intensity of the sound contained in the audio frame is within the sound intensity range corresponding to the specified object, then generating an analysis result representing that the audio frame contains sound information; if the frequency of the sound contained in the audio frame is not within the sound frequency range corresponding to the specified object, or the intensity of the sound contained in the audio frame is not within the sound intensity range corresponding to the specified object, then generating an analysis result representing that the audio frame does not contain sound information.
[0067] In one embodiment, when the sound information is a sound containing specified content, the specific process of audio analysis may include: performing speech recognition processing on the sound contained in the audio frame to obtain the audio text corresponding to the audio frame; obtaining the text matching degree between the audio text and the specified content; if the text matching degree is greater than or equal to the preset matching degree, then generating an analysis result indicating that the audio frame contains sound information; if the text matching degree is less than the preset matching degree, then generating an analysis result indicating that the audio frame does not contain sound information. The preset matching degree may be set manually or by the server or terminal device in the above-mentioned audio data processing system, which is not limited here. Optionally, audio analysis may also be performed in other ways, which are not limited here.
[0068] S203: If the analysis result indicates that the audio frame does not contain sound information, the audio frame is discarded, and an audio recording file corresponding to the voice conversation is generated based on the remaining audio frames in the conversation audio data.
[0069] In an embodiment of the present application, the discarding process may include deleting audio frames, transferring audio frames to other storage spaces, stopping audio recording, and any other operations that do not record audio frames, which are not limited here.
[0070] Optionally, an audio frame contains a short audio segment, such as 10ms or 20ms. People may pause briefly when speaking, and animals may take short breaks when roaring. Although these short pauses do not contain valid sound information, they make it easier for the recorder to distinguish and understand the valid information in the audio, thereby improving the recorder's user experience.
[0071] Therefore, if the analysis result indicates that the audio frame does not contain sound information, the specific process of discarding the audio frame may include: if the analysis results corresponding to adjacent audio frames all indicate that the audio frame does not contain sound information, and the total audio duration corresponding to the adjacent audio frames is greater than a preset duration, then discarding the adjacent audio frames. Correspondingly, if the analysis results corresponding to adjacent audio frames all indicate that the audio frame does not contain sound information, and the total audio duration corresponding to the adjacent audio frames is less than or equal to the preset duration, then recording the adjacent audio frames.
[0072] The number of adjacent audio frames can include two or more, and adjacent audio frames refer to consecutive audio frames generated during a voice conversation. The total audio duration corresponding to the adjacent audio frames is calculated by adding the audio duration corresponding to each audio frame in the adjacent audio frames. Furthermore, since the aforementioned voice conversation continuously generates conversation audio data, adjacent audio frames may belong to the same conversation audio data or may be from two consecutive conversation audio data.
[0073] The preset duration can be set manually or by the server or terminal device in the audio data processing system, which is not limited here. Specifically, the preset duration can be set based on the recording person's listening habits.
[0074] Furthermore, the analysis results corresponding to the remaining audio frames in the conversation audio data indicate that the remaining audio frames contain sound information. The specific process of generating an audio recording file corresponding to the voice conversation based on the remaining audio frames in the conversation audio data may include: recording the remaining audio frames in the conversation audio data to obtain the audio recording file. The recording process may include audio recording, file encoding, and file compression.
[0075] In one embodiment, since the voice conversation mentioned in step S201 will continuously generate conversation audio data, the specific process of generating an audio recording file corresponding to the voice conversation based on the remaining audio frames in the conversation audio data may include: recording the remaining audio frames in the conversation audio data, and triggering the step of obtaining the generated conversation audio data until the voice conversation ends, thereby generating an audio recording file corresponding to the voice conversation.
[0076] For specific implementation, please refer to the attached Figure 3, shows a schematic diagram of an audio recording process. Figure 3 As shown, the local client can collect local audio data through the audio collection thread during the voice conversation; then, the local client can perform audio quality enhancement on the collected local audio data and write the processed local audio data into the audio cache thread; at the same time, the local client will also encode the processed local audio data and upload the encoded data to the server.
[0077] In addition, if Figure 3 As shown, the local client also synchronously receives and caches the remote audio data collected by the remote client sent by the server. After decoding the remote audio data, the local client plays the remote audio data through the audio playback thread and writes the remote audio data to the audio cache thread.
[0078] The audio cache thread mixes the written local audio data and remote audio data to obtain the conversation audio data of the voice conversation, and sends the conversation audio data to the audio analysis thread, which performs audio analysis on each audio frame in the conversation audio data. If the analysis result of a certain audio frame indicates that the audio frame contains sound information, the audio analysis thread will start the recording function of the audio recording thread, and the audio recording thread will record the audio frame; if the analysis results of multiple consecutive audio frames all indicate that the audio frame does not contain sound information, and the audio duration corresponding to the multiple audio frames is greater than the preset duration, the audio analysis thread will turn off the recording function of the audio recording thread, thereby discarding the multiple audio frames.
[0079] In an embodiment of the present application, by performing audio analysis on the audio frames contained in the conversation audio data, it is possible to accurately identify which audio frames do not contain sound information and are therefore discarded, and which audio frames contain sound information and are therefore recorded; the audio recording files obtained in this way all contain sound information, and sound information can convey effective information in the voice conversation, thereby greatly improving the proportion of effective information in the audio recording files.
[0080] In addition, the embodiment of the present application obtains the conversation audio data generated during the voice conversation and performs audio analysis synchronously. Compared with the method of first recording all the conversation audio data generated during the voice conversation, then performing audio analysis and editing the audio based on the analysis results; the embodiment of the present application can reuse the time during the voice conversation, so that the recorder can obtain an audio recording file with a high proportion of valid information in a short time after the end of the voice conversation, which is beneficial to improving the recorder's user experience. At the same time, the embodiment of the present application will promptly discard audio frames that do not contain valid information during the voice conversation, avoiding audio frames with invalid information from occupying more resources, which is beneficial to improving the processing efficiency of audio data.
[0081] In one embodiment of the present application, another audio data processing method is provided, which can be applied to Figure 1 The audio data processing system shown in FIG. 1 can be executed by a terminal device or a server, or can be executed by both a terminal device and a server. In the embodiment of the present application, the method is described by taking the method executed by a terminal device as an example. Figure 4 FIG. 1 is a flow chart showing another method for processing audio data, wherein the method is Figure 2 This paper expands on the method shown in .
[0082] Among them, S401 to S405 are described in detail as follows:
[0083] S401: During a voice conversation, obtain generated conversation audio data.
[0084] In the embodiment of the present application, the specific implementation of step S401 can refer to the specific implementation of step S201 in the above embodiment, which will not be repeated here.
[0085] S402: Perform audio perception detection on each audio frame in the conversation audio data to obtain a perception result.
[0086] In an embodiment of the present application, audio perception detection is used to detect whether an audio frame contains audio information within the perception range of a preset object. The preset object and the specified object may be the same or different. The preset object may be specifically set based on the needs of the recorder. For example, the voice conversation may be an animal live broadcast, and the purpose of the recorder recording the animal live broadcast is to subsequently drive away another animal in the wild, then the preset object may be another animal. For another example, the voice conversation may be an online meeting, and the recorder records the online meeting for the purpose of subsequently taking meeting minutes, then the preset object may be a person.
[0087] In one embodiment, audio perception detection mainly determines whether the audio energy contained in the audio frame can reach the energy range of the sound that can be heard by the preset object; the audio collection data is a 2-byte integer with a maximum value of 32767 and a minimum value of 0; the audio energy of an audio frame is the maximum data amount of the data segment in the audio frame.
[0088] Therefore, the specific process of audio perception detection may include: obtaining the data volume of each data segment in the audio frame; if the maximum data volume among the obtained data volumes is greater than or equal to a preset data volume threshold, generating a perception result representing audio information contained in the audio frame within the perception range of a preset subject; if the maximum data volume among the obtained data volumes is less than the preset data volume threshold, generating a perception result representing audio information not contained in the audio frame within the perception range of the preset subject. The preset data volume threshold may be set based on the energy range of sound audible to the preset subject.
[0089] For example, the preset object can be set to be a person, and the preset data volume threshold can be 5 bits. Figure 5 , shows a schematic diagram of the data volume of a data slice. Figure 5 As shown, the data amount of each data slice in the audio frame 501 is 16 bits; then, the maximum data amount in the audio frame 501 is 16 bits, which is greater than 5 bits; therefore, a perception result representing the audio information contained in the audio frame within the perception range of the human ear can be generated.
[0090] S403: If the perception result indicates that the audio frame contains audio information within the perception range of the preset object, perform sound detection on the audio frame with respect to the specified object to obtain a sound detection result, and generate an analysis result based on the sound detection result.
[0091] In an embodiment of the present application, the sound detection process for a specified object may specifically include: performing data conversion processing on an audio frame to obtain a frequency domain signal corresponding to the audio frame; obtaining the frequency domain energy of a signal segment within a preset frequency range in the frequency domain signal; if the ratio between the obtained frequency domain energy and the total frequency domain energy of the frequency domain signal is greater than the preset ratio, generating a sound detection result representing that the audio frame contains sound information of the specified object; if the ratio between the obtained frequency domain energy and the total frequency domain energy of the frequency domain signal is less than or equal to the preset ratio, generating a sound detection result representing that the audio frame does not contain sound information of the specified object.
[0092] The preset frequency range is set based on the sound frequency range of the specified object. The preset ratio can be set manually or by the server or terminal device in the above audio data processing system, which is not limited here.
[0093] Specifically, the frequency domain energy of the frequency domain signal and the frequency domain energy of the signal segment within the preset frequency range can be calculated using the frequency domain energy calculation formula. Since the frequency domain energy calculation formula is well known to those skilled in the art, it will not be described in detail here.
[0094] Furthermore, the specific process of generating analysis results based on sound detection results includes: if the sound detection result indicates that the audio frame contains sound information of a specified object, then generating an analysis result indicating that the audio frame contains sound information; if the sound detection result indicates that the audio frame does not contain sound information of the specified object, then generating an analysis result indicating that the audio frame contains sound information.
[0095] In the specific implementation, the specified object can be set as a person, and the sound detection of the specified object can be specifically voice detection (VoiceActivityDetection, VAD, used to identify whether the audio contains human voice). Figure 6 , shows a schematic diagram of a VAD detection process. Figure 6 As shown, after the audio frame is input into the VAD detection module, the VAD detection module processes the audio frame to obtain and output a human voice detection result; wherein the human voice detection result is used to indicate whether the audio frame contains human voice.
[0096] The specific processing of the audio frame by the VAD detection module may include: performing data conversion processing on the audio frame to obtain the frequency domain signal corresponding to the audio frame; dividing the frequency domain signal into one or more signal segments based on a preset cycle duration, with the duration of each signal segment being the same as the preset cycle duration. Obtaining the human voice frequency domain energy of each signal segment within the human voice frequency range and the total frequency domain energy of each signal segment; if the ratio between the human voice frequency domain energy corresponding to any signal segment and the total frequency domain energy corresponding to any signal segment is greater than a preset ratio, generating a human voice detection result indicating that the audio frame contains human voice; if the ratio between the human voice frequency domain energy corresponding to each signal segment in adjacent signal segments and the total frequency domain energy corresponding to each signal segment in adjacent signal segments is less than or equal to the preset ratio, generating a human voice detection result indicating that the audio frame does not contain human voice. The preset cycle duration and the number of adjacent signal segments can be manually set or can be set by the server or terminal device in the above-mentioned audio data processing system, and are not limited here. Specifically, the preset cycle duration can be 10ms, 20ms, etc. The number of adjacent signal segments can be specifically set based on human listening habits and the duration of speaking pauses.
[0097] S404: If the perception result indicates that the audio frame does not contain audio information within the perception range of the preset object, generate an analysis result indicating that the audio frame does not contain sound information.
[0098] It can be seen from steps S402 and S403 that the amount of computation and resources required for audio perception detection is much smaller than the amount of computation and resources required for sound detection. Therefore, the embodiment of the present application can perform preliminary screening of audio frames before sound detection by first performing audio perception detection and then performing sound detection, thereby reducing the number of audio frames that need to be subsequently sound detected, which is beneficial to improving the efficiency of audio analysis of audio frames and also reduces the consumption of computing resources.
[0099] In one possible implementation, in scenarios such as online meetings and interactive live broadcasts, music can convey emotions and atmosphere, and also provide useful information for the recorder. Therefore, sound information can also include music information.
[0100] Then, the specific process of generating analysis results based on sound detection results may include: if the sound detection result indicates that the audio frame does not contain sound information of the specified object, then the audio frame can be further subjected to music detection to obtain a music detection result; if the music detection result indicates that the audio frame contains music information, then an analysis result indicating that the audio frame contains sound information can be generated; if the music detection result indicates that the audio frame does not contain music information, then an analysis result indicating that the audio frame does not contain sound information is generated.
[0101] Optionally, the specific process of music detection may include: if the frequency of the sound contained in the audio frame is within the music frequency range, then a music detection result indicating that the audio frame contains music information may be generated; if the frequency of the sound contained in the audio frame is not within the music frequency range, then a music detection result indicating that the audio frame does not contain music information may be generated. In practical applications, the music frequency range is approximately 20 Hz to 20 kHz.
[0102] Optionally, the specific process of music detection may also include: performing data conversion processing on the audio frame to obtain the frequency domain signal corresponding to the audio frame; obtaining the frequency domain energy of the signal segment within the music frequency range in the frequency domain signal; if the obtained frequency domain energy is greater than a preset energy threshold, then generating a music detection result indicating that the audio frame contains music information; if the obtained frequency domain energy is less than or equal to the preset energy threshold, then generating a music detection result indicating that the audio frame does not contain music information. The preset energy threshold may be set manually or by a server or terminal device in the above-mentioned audio data processing system, which is not limited here. Specifically, the preset energy threshold may be set based on the maximum audio energy and minimum audio energy that the music can generate.
[0103] Please see the attached Figure 7 , shows a schematic diagram of a music detection process. Figure 7As shown, after the audio frame is input into the music detection module, the music detection module processes the audio frame to obtain and output a music detection result; wherein the music detection result is used to indicate whether the audio frame contains music information.
[0104] The specific processing of the audio frame by the music detection module may include: performing data conversion processing on the audio frame to obtain the frequency domain signal corresponding to the audio frame; dividing the frequency domain signal into one or more signal segments based on the specified cycle duration, and the duration of each signal segment is the same as the specified cycle duration. Obtaining the music frequency domain energy of each signal segment within the music frequency range; if the music frequency domain energy corresponding to any signal segment is greater than the preset energy threshold, then generating a music detection result indicating that the audio frame contains music; if the music frequency domain energy corresponding to each signal segment in the continuous signal segments is less than the preset energy threshold, then generating a music detection result indicating that the audio frame does not contain music. Among them, the specified cycle duration and the number of continuous signal segments can be set manually, or can be set by the server or terminal device in the above-mentioned audio data processing system, and are not limited here. Specifically, the specified cycle duration can be 10ms, 20ms, etc. The number of continuous signal segments can be specifically set based on the pause duration of the music segment.
[0105] S405: If the analysis result indicates that the audio frame does not contain sound information, the audio frame is discarded, and an audio recording file corresponding to the voice conversation is generated based on the remaining audio frames in the conversation audio data.
[0106] In the embodiment of the present application, the specific implementation of step S405 can refer to the specific implementation of step S203 in the above embodiment, which will not be repeated here.
[0107] In actual application, both the designated object and the preset object can be set to be a person. Figure 8 , shows a schematic diagram of the audio analysis process. Figure 8 As shown, during the audio analysis of conversation audio data, audio perception detection can be performed on the audio frames according to the order in which the audio frames are arranged in the conversation audio data. If a sound is detected in the audio frame that is not within the human ear's perception range, it can be determined that the audio frame does not contain valid information, thereby generating an analysis result indicating that no sound information is contained. If a sound is detected in the audio frame that is within the human ear's perception range, it can be determined that the audio frame may contain a human voice, and further VAD detection can be performed.
[0108] If VAD detection finds a human voice in an audio frame, it can be determined that the audio frame contains valid information, generating an analysis result indicating that it contains sound information. Otherwise, music detection is required. If music detection finds music in an audio frame, it can be determined that the audio frame contains valid information, generating an analysis result indicating that it contains sound information. Otherwise, it can be determined that the audio frame does not contain valid information, generating an analysis result indicating that it does not contain sound information.
[0109] In an embodiment of the present application, by performing sound detection on an audio frame regarding a specified object, it is possible to accurately identify whether the audio frame contains sound information of the specified object, which is conducive to the subsequent generation of more accurate analysis results, thereby increasing the proportion of effective information in the audio recording file. At the same time, the embodiment of the present application performs an audio perception detection method that requires less computing resources before sound detection, and can perform a preliminary screening of the audio frames before sound detection, thereby reducing the number of audio frames that need to be sound detected later, which is conducive to improving the efficiency of audio analysis of audio frames and reducing computing resource consumption.
[0110] Here, we introduce the device embodiment of the present application, which can be used to execute the audio data processing method in the above embodiment of the present application. For details not disclosed in the device embodiment of the present application, please refer to the embodiment of the audio data processing method in the above embodiment of the present application.
[0111] The present application provides an audio data processing device, such as Figure 9 As shown, the apparatus includes an acquisition unit 901 and a processing unit 902, wherein:
[0112] The acquisition unit 901 is used to acquire the conversation audio data generated during the voice conversation;
[0113] The processing unit 902 is configured to perform audio analysis on each audio frame in the conversation audio data to obtain an analysis result;
[0114] The processing unit 902 is further configured to discard the audio frame if the analysis result indicates that the audio frame does not contain sound information, and generate an audio recording file corresponding to the voice conversation based on the remaining audio frames in the conversation audio data.
[0115] In one embodiment of the present application, conversation audio data is continuously generated during a voice conversation. Based on the aforementioned scheme, when the processing unit 902 generates an audio recording file corresponding to the voice conversation based on the remaining audio frames in the conversation audio data, it can be specifically used to: record the remaining audio frames in the conversation audio data, and trigger the step of obtaining the generated conversation audio data until the voice conversation ends, thereby generating an audio recording file corresponding to the voice conversation.
[0116] In one embodiment of the present application, based on the aforementioned scheme, when the processing unit 902 discards the audio frame if the analysis result indicates that the audio frame does not contain sound information, it can be specifically used to: if the analysis results corresponding to adjacent audio frames both indicate that the audio frame does not contain sound information, and the total audio duration corresponding to the adjacent audio frames is greater than the preset duration, then the adjacent audio frames are discarded.
[0117] In one embodiment of the present application, the sound information includes sound information of a specified object; based on the aforementioned scheme, when the processing unit 902 performs audio analysis on the audio frame to obtain an analysis result, it can be specifically used to: perform audio perception detection on the audio frame to obtain a perception result; if the perception result represents that the audio frame contains audio information within the perception range of a preset object, then perform sound detection on the audio frame regarding the specified object to obtain a sound detection result, and generate an analysis result based on the sound detection result; if the perception result represents that the audio frame does not contain audio information within the perception range of the preset object, then generate an analysis result representing that the audio frame does not contain sound information.
[0118] In one embodiment of the present application, based on the aforementioned scheme, when the processing unit 902 performs sound detection on the audio frame regarding the specified object and obtains the sound detection result, it can be specifically used to: perform data conversion processing on the audio frame to obtain the frequency domain signal corresponding to the audio frame; obtain the frequency domain energy of the signal segment within the preset frequency range in the frequency domain signal; the preset frequency range is set based on the sound frequency range of the specified object; if the ratio between the obtained frequency domain energy and the total frequency domain energy of the frequency domain signal is greater than the preset ratio, a sound detection result is generated representing that the audio frame contains sound information of the specified object; if the ratio between the obtained frequency domain energy and the total frequency domain energy of the frequency domain signal is less than or equal to the preset ratio, a sound detection result is generated representing that the audio frame does not contain sound information of the specified object.
[0119] In one embodiment of the present application, the sound information also includes music information. Based on the aforementioned scheme, when the processing unit 902 generates an analysis result based on the sound detection result, it can be specifically used to: if the sound detection result indicates that the audio frame does not contain the sound information of the specified object, then perform music detection on the audio frame to obtain a music detection result; if the music detection result indicates that the audio frame contains music information, then generate an analysis result indicating that the audio frame contains sound information; if the music detection result indicates that the audio frame does not contain music information, then generate an analysis result indicating that the audio frame does not contain sound information.
[0120] In one embodiment of the present application, based on the aforementioned scheme, when the processing unit 902 performs music detection on the audio frame and obtains the music detection result, it can also be used to: perform data conversion processing on the audio frame to obtain the frequency domain signal corresponding to the audio frame; obtain the frequency domain energy of the signal segment within the music frequency range in the frequency domain signal; if the obtained frequency domain energy is greater than a preset energy threshold, a music detection result is generated indicating that the audio frame contains music information; if the obtained frequency domain energy is less than or equal to the preset energy threshold, a music detection result is generated indicating that the audio frame does not contain music information.
[0121] In one embodiment of the present application, the device is configured on a local client. Based on the aforementioned scheme, when the processing unit 902 obtains the session audio data generated during the voice conversation, it can be specifically used to: collect local audio data generated by the local client in the voice conversation during the voice conversation; obtain remote audio data generated by the remote client joining the voice conversation in the voice conversation; mix the local audio data and the remote audio data to obtain the session audio data of the voice conversation.
[0122] In one embodiment of the present application, based on the aforementioned scheme, the processing unit 902 can also be used to: perform audio quality enhancement processing on the audio frames contained in the local audio data according to the arrangement order of the audio frames in the local audio data to obtain enhanced audio frames; encode the enhanced audio frames; and whenever an encoded enhanced audio frame is detected, send the encoded enhanced audio frame to the server, so that the server sends the encoded enhanced audio frame to the remote client whenever it receives the encoded enhanced audio frame.
[0123] It should be noted that the apparatus provided in the above embodiment and the method provided in the above embodiment belong to the same concept, wherein the specific manner in which each module and unit performs operations has been described in detail in the method embodiment and will not be repeated here.
[0124] The device provided in the above embodiment can be arranged in a terminal device or in a server. The device provided in the embodiment of the present application can accurately identify which audio frames do not contain sound information and are discarded, and which audio frames contain sound information and are recorded by performing audio analysis on the audio frames contained in the conversation audio data; the audio recording files obtained in this way all contain sound information, and the sound information can convey effective information in the voice conversation, thereby greatly improving the proportion of effective information in the audio recording files.
[0125] An embodiment of the present application also provides an electronic device, comprising one or more processors and a storage device, wherein the storage device is used to store one or more computer programs, and when the one or more computer programs are executed by one or more processors, the electronic device implements the above audio data processing method.
[0126] Figure 10 A schematic diagram of the structure of a computer system suitable for implementing an electronic device according to an embodiment of the present application is shown.
[0127] It should be noted that Figure 10 The computer system 1000 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0128] like Figure 10 As shown, the computer system 1000 includes a processor (Central Processing Unit, CPU) 1001, which can perform various appropriate actions and processes according to the program stored in the read-only memory (Read-Only Memory, ROM) 1002 or the program loaded from the storage part 1008 into the random access memory (Random Access Memory, RAM) 1003, such as executing the method in the above embodiment. Various programs and data required for system operation are also stored in the RAM 1003. The CPU 1001, ROM 1002 and RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0129] In some embodiments, the following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, and the like; an output section 1007 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 1008 including a hard disk; and a communication section 1009 including a network interface card such as a LAN (Local Area Network) card or a modem. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. Removable media 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like, is installed in the drive 1010 as needed, so that computer programs read therefrom can be installed into the storage section 1008 as needed.
[0130] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1009, and / or installed from a removable medium 1011. When the computer program is executed by the processor (CPU) 1001, the various functions defined in the system of the present application are executed.
[0131] It should be noted that the computer-readable medium shown in the embodiment of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (Erasable Programmable Read Only Memory), a flash memory, an optical fiber, a portable compact disk read-only memory (Compact Disc Read-Only Memory, CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, wherein a computer-readable computer program is carried. This propagated data signal can take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. A computer program embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0132] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions and operations of the devices, methods and computer program products according to various embodiments of the present application. Among them, each box in the flowchart or block diagram can represent a module, program segment, or part of the code, and the above-mentioned module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and a computer program.
[0133] The units or modules described in the embodiments of the present application may be implemented in software or hardware, and the units or modules described may also be provided in a processor. The names of these units or modules do not, in certain circumstances, limit the units or modules themselves.
[0134] Another aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the aforementioned audio data processing method. The computer-readable storage medium may be included in the electronic device described in the above embodiments, or may exist independently and not be incorporated into the electronic device.
[0135] Another aspect of the present application further provides a computer program product, comprising a computer program stored in a computer-readable storage medium. A processor of an electronic device reads the computer program from the computer-readable storage medium and executes the computer program, causing the electronic device to perform the audio data processing method described above in each of the above embodiments.
[0136] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiment of the application, the features and functions of two or more modules or units described above can be concretized in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.
[0137] Other embodiments of the present invention will readily occur to those skilled in the art after considering the specification and practicing the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of this application and include common knowledge or customary techniques in the art that are not disclosed herein.
[0138] The above content is only a preferred exemplary embodiment of the present application and is not intended to limit the implementation scheme of the present application. Ordinary technicians in this field can easily make corresponding changes or modifications based on the main concept and spirit of the present application. Therefore, the scope of protection of the present application shall be based on the scope of protection required by the claims.
Claims
1. A method for processing audio data, characterized in that: The method comprises: During the voice conversation, obtaining the generated conversation audio data; For each audio frame in the conversation audio data, performing audio analysis on the audio frame to obtain an analysis result; If the analysis result indicates that the audio frame does not contain sound information, the audio frame is discarded, and an audio recording file corresponding to the voice conversation is generated based on the remaining audio frames in the conversation audio data.
2. The method according to claim 1, characterized in that During the voice conversation, conversation audio data is continuously generated; and generating an audio recording file corresponding to the voice conversation based on remaining audio frames in the conversation audio data includes: The remaining audio frames in the conversation audio data are recorded and the step of obtaining the generated conversation audio data is triggered until the voice conversation ends, thereby generating an audio recording file corresponding to the voice conversation.
3. The method according to claim 1, characterized in that If the analysis result indicates that the audio frame does not contain sound information, discarding the audio frame includes: If the analysis results corresponding to adjacent audio frames both indicate that the audio frames do not contain sound information, and the total audio duration corresponding to the adjacent audio frames is greater than a preset duration, the adjacent audio frames are discarded.
4. The method according to claim 1, wherein The sound information includes sound information of a specified object; and the performing of audio analysis on the audio frame to obtain an analysis result includes: Performing audio perception detection on the audio frame to obtain a perception result; If the perception result indicates that the audio frame contains audio information within the perception range of a preset object, performing sound detection on the audio frame with respect to the specified object to obtain a sound detection result, and generating the analysis result based on the sound detection result; If the perception result indicates that the audio frame does not contain audio information within the perception range of the preset object, an analysis result indicating that the audio frame does not contain sound information is generated.
5. The method according to claim 4, characterized in that The performing sound detection on the audio frame about the specified object to obtain a sound detection result includes: Performing data conversion processing on the audio frame to obtain a frequency domain signal corresponding to the audio frame; Acquiring frequency domain energy of a signal segment within a preset frequency range in the frequency domain signal; the preset frequency range is set based on a sound frequency range of the designated object; If a ratio between the acquired frequency domain energy and the total frequency domain energy of the frequency domain signal is greater than a preset ratio, generating a sound detection result indicating that the audio frame contains sound information of the specified object; If the ratio of the acquired frequency domain energy to the total frequency domain energy of the frequency domain signal is less than or equal to a preset ratio, a sound detection result is generated indicating that the audio frame does not contain sound information of the specified object.
6. The method according to claim 4, characterized in that The sound information also includes music information, and generating the analysis result based on the sound detection result includes: If the sound detection result indicates that the audio frame does not contain the sound information of the specified object, performing music detection on the audio frame to obtain a music detection result; If the music detection result indicates that the audio frame contains music information, generating an analysis result indicating that the audio frame contains sound information; If the music detection result indicates that the audio frame does not contain music information, an analysis result indicating that the audio frame does not contain sound information is generated.
7. The method according to claim 6, characterized in that The performing music detection on the audio frame to obtain a music detection result includes: Performing data conversion processing on the audio frame to obtain a frequency domain signal corresponding to the audio frame; Acquiring frequency domain energy of a signal segment within a music frequency range in the frequency domain signal; If the acquired frequency domain energy is greater than a preset energy threshold, a music detection result is generated indicating that the audio frame contains music information; If the acquired frequency domain energy is less than or equal to the preset energy threshold, a music detection result is generated indicating that the audio frame does not contain music information.
8. The method according to any one of claims 1 to 7, characterized in that The method is applied to a local client, and obtaining the generated conversation audio data during the voice conversation process includes: During the voice conversation, collecting local audio data generated by the local client in the voice conversation; Acquire remote audio data generated by a remote client joining the voice session in the voice session; The local audio data and the remote audio data are mixed to obtain session audio data of the voice session.
9. The method according to claim 8, characterized in that The method further comprises: performing audio quality enhancement processing on the audio frames contained in the local audio data according to an arrangement order of the audio frames in the local audio data to obtain enhanced audio frames; encoding the enhanced audio frame; Whenever an encoded enhanced audio frame is detected, the encoded enhanced audio frame is sent to the server, so that whenever the server receives the encoded enhanced audio frame, it sends the encoded enhanced audio frame to the remote client.
10. An audio data processing device, characterized in that: The device comprises an acquisition unit and a processing unit, wherein: The acquisition unit is used to acquire the conversation audio data generated during the voice conversation; The processing unit is configured to perform audio analysis on each audio frame in the conversation audio data to obtain an analysis result; The processing unit is further configured to discard the audio frame if the analysis result indicates that the audio frame does not contain sound information, and generate an audio recording file corresponding to the voice conversation based on the remaining audio frames in the conversation audio data.
11. A computer-readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the audio data processing method according to any one of claims 1 to 9 is implemented.
12. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the one or more processors to implement the audio data processing method according to any one of claims 1 to 9.