Audio classification method, device and computer-readable storage medium

By converting audio into audio text and classifying it in combination with text reference database, the problem of low audio classification efficiency is solved, and the audio information of specified content is quickly obtained.

CN113590871BActive Publication Date: 2025-08-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110163903.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-05
Publication Date
2025-08-29
Estimated Expiration
2041-02-05

AI Technical Summary

Technical Problem

In the prior art, audio classification is inefficient and cannot quickly obtain specified content, resulting in large time loss when listening to audio information.

Method used

By converting audio into audio text and using a text reference database for classification, identifying and highlighting the target text content, and determining the classification result in response to user operations.

Benefits of technology

It improves the efficiency of audio classification, reduces the fatigue of classifiers, and reduces the time loss of listening to audio information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113590871B_ABST
    Figure CN113590871B_ABST
Patent Text Reader

Abstract

The present invention discloses an audio classification method, device, and computer-readable storage medium. Embodiments of the present application can display an audio classification page, which includes an audio text converted from the audio to be classified, and a classification control for the audio text. The audio text includes a highlighted target text content, which is text content identified from the audio text that matches preset text content in a text reference database. Each classification control corresponds to a classification result. The conversion between the audio to be classified and the audio text can be implemented based on voice technology in the field of artificial intelligence. In response to a classification operation on the classification control, the classification control operated by the classification operation is determined to be a target classification control, the classification result corresponding to the target classification control in the target text content is determined, and the classification result of the audio to be classified is determined based on the classification result of the target text content. This solution can improve the efficiency of audio classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of communication technology, and in particular to an audio classification method, device, and computer-readable storage medium. Background Art

[0002] With the increase in Internet users, the China Internet Network Information Center (CNNIC) released the "26th Statistical Report on the Development of China's Internet" (hereinafter referred to as the "Report"). The report shows that the number of Internet users in China has reached 420 million, and the number of mobile Internet users has reached 277 million. The amount of information on various Internet websites is huge.

[0003] During the research and practice of relevant technologies, the inventors of this application found that for websites with a large amount of information, the form of network information includes pictures, text, audio and other content. Relatively speaking, pictures and text can be simply read by the naked eye, but audio information needs to be completed by "listening", and it is difficult to do actions such as fast forward and fast rewind like songs. It needs to be listened to from beginning to end, which results in a large loss of time invested in audio listening. It is impossible to quickly obtain specified content while listening to audio information content, and the efficiency of audio classification is low. Summary of the Invention

[0004] The embodiments of the present application provide an audio classification method, apparatus, and computer-readable storage medium, which can improve the efficiency of audio classification.

[0005] This embodiment of the present application provides an audio classification method, including:

[0006] Displaying an audio classification page, the audio classification page including an audio text converted from the audio to be classified and a classification control for the audio text, wherein the audio text includes a highlighted target text content, the target text content being text content identified from the audio text that matches preset text content in a text reference database, and one classification control corresponding to one classification result;

[0007] In response to a classification operation on a target classification control, the classification control operated by the classification operation is determined to be a target classification control, the classification result corresponding to the target classification control in the target text content is determined, and based on the classification result of the target text content, the classification result of the audio to be classified is determined.

[0008] Accordingly, an embodiment of the present application provides an audio classification device, comprising:

[0009] a page display unit, configured to display an audio classification page, the audio classification page including an audio text converted from the audio to be classified, and a classification control for the audio text, wherein the audio text includes a highlighted target text content, the target text content being text content identified from the audio text that matches a preset text content in a text reference database, and one classification control corresponding to one classification result;

[0010] A result determination unit is used to respond to a classification operation on a classification control, determine that the classification control operated by the classification operation is a target classification control, determine the classification result corresponding to the target classification control in the target text content, and determine the classification result of the audio to be classified based on the classification result of the target text content.

[0011] In one embodiment, the page display unit includes:

[0012] a receiving subunit, configured to receive an audio classification request for audio to be classified, and obtain the audio to be classified based on the audio classification request;

[0013] A first recognition subunit is configured to perform content recognition on the audio to be classified, and convert the audio to be classified into audio text based on the content recognition result;

[0014] The first page display subunit is used to display the audio classification page based on the audio text.

[0015] In one embodiment, the page display unit includes:

[0016] A second recognition subunit is configured to recognize the audio text to be classified of the audio to be classified based on a text reference database;

[0017] a changing subunit, configured to change the display form of the target text content to highlighting when identifying that the target text content includes the preset text content in the text reference database in the audio text to be classified;

[0018] The second page display subunit is used to display the audio classification page based on the highlighting result of the target text content.

[0019] In one embodiment, the audio classification device further includes:

[0020] The first playback unit is used to play the target audio corresponding to the target text content in the audio to be classified in response to a trigger operation on the target text content when the classification result of the audio to be classified is failure, so as to verify the classification result of the audio to be classified.

[0021] In one embodiment, the first playback unit includes:

[0022] an information determination subunit, configured to, when a classification result of the audio to be classified is a failure, determine, in response to a triggering operation on the target text content, time information of the target audio corresponding to the target text content within the audio to be classified;

[0023] The playing subunit is configured to play the target audio corresponding to the time information in response to a triggering operation on the audio playing control, so as to verify the classification result of the audio to be classified.

[0024] In one embodiment, the audio classification device further includes:

[0025] A changing unit is used to change the classification result of the audio to be classified in response to a switching operation on other classification controls when the playback result of the target audio does not match the target text content. The other classification controls are controls in the classification controls other than the target classification control.

[0026] In one embodiment, the audio classification device further includes:

[0027] an acquiring unit, configured to acquire additional information corresponding to a changed classification result of the audio to be classified, wherein the additional information includes description information of the changed classification result;

[0028] A sending unit is configured to send a classification result of the audio to be classified to an initiating terminal of the audio to be classified based on the description information.

[0029] In one embodiment, the audio classification device further includes:

[0030] a second playing unit configured to, when a classification result of the audio to be classified is a failure, respond to a trigger operation on a subtext content, determine that the subtext content corresponding to the trigger operation is a target subtext content, and play a target sub-audio corresponding to the target subtext content in the audio to be classified;

[0031] The third playback unit is used to play the sub-audio corresponding to the other sub-text content in the audio to be classified in response to a trigger operation for other sub-text content when the playback result of the target sub-audio does not match the target sub-text content, so as to verify the classification result of the audio to be classified. The other sub-text content is the text content in the multiple sub-text contents except the target sub-text content.

[0032] Accordingly, an embodiment of the present application also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the steps of any audio classification method provided in the embodiment of the present application.

[0033] Accordingly, an embodiment of the present application further provides a computer-readable storage medium, which stores a plurality of instructions, and the instructions are suitable for loading by a processor to execute the steps in any audio classification method provided in the embodiment of the present application.

[0034] An embodiment of the present application can display an audio classification page, which includes an audio text after the audio to be classified is converted, and a classification control for the audio text, wherein the audio text includes a highlighted target text content, and the target text content is text content identified from the audio text that matches the preset text content in the text reference database, and one classification control corresponds to one classification result; in response to a classification operation on the classification control, the classification control operated by the classification operation is determined to be a target classification control, the classification result corresponding to the target classification control in the target text content is determined, and the classification result of the audio to be classified is determined based on the classification result of the target text content. This solution can convert the audio to be classified into an audio text, and by converting the audio to be classified into an audio text and then classifying the audio text to determine the classification result of the audio to be classified, the fatigue of the classification personnel who play the audio to be classified audio in a loop and judge the audio content can be alleviated, and the efficiency of audio classification can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0036] Figure 1 This is a schematic diagram of a scenario of the audio classification method provided in an embodiment of the present application;

[0037] Figure 2a This is a flowchart of the audio classification method provided by an embodiment of the present application;

[0038] Figure 2b This is a schematic diagram of an audio classification page of the audio classification method provided in an embodiment of the present application;

[0039] Figure 2c This is another schematic diagram of an audio classification page of the audio classification method provided in an embodiment of the present application;

[0040] Figure 2d This is another schematic diagram of an audio classification page of the audio classification method provided in an embodiment of the present application;

[0041] Figure 3a is another flow chart of the audio classification method provided in an embodiment of the present application;

[0042] Figure 3b This is a technical flow chart of the audio classification method provided by an embodiment of the present application;

[0043] Figure 3c This is another schematic diagram of an audio classification page of the audio classification method provided in an embodiment of the present application;

[0044] Figure 4a is a device diagram of the audio classification method provided in an embodiment of the present application;

[0045] Figure 4b is another apparatus diagram of the audio classification method provided in an embodiment of the present application;

[0046] Figure 4c is another apparatus diagram of the audio classification method provided in an embodiment of the present application;

[0047] Figure 5 It is a structural diagram of the computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0048] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.

[0049] The embodiments of the present application provide an audio classification method, apparatus, computer device, and computer-readable storage medium. Specifically, the embodiments of the present application provide an audio classification device suitable for a computer device. The computer device may be a terminal or a server, and the server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. The terminal may be a smartphone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server may be directly or indirectly connected via wired or wireless communication, and this application does not impose any restrictions thereon.

[0050] refer to Figure 1Taking a computer device as a terminal as an example, the terminal can display an audio classification page, which includes an audio text converted from the audio to be classified, and a classification control of the audio text, wherein the audio text includes a highlighted target text content, and the target text content is text content identified from the audio text that matches the preset text content in the text reference database, and one classification control corresponds to one classification result; in response to a classification operation on the classification control, the classification control operated by the classification operation is determined to be a target classification control, the classification result corresponding to the target classification control in the target text content is determined, and based on the classification result of the target text content, the classification result of the audio to be classified is determined.

[0051] Among them, converting the audio to be classified into audio text can be achieved based on voice technology in the field of artificial intelligence. For example, the audio content of the audio to be classified can be identified through voice technology, and then the identified audio content can be converted into audio text.

[0052] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or computer-controlled machine models to extend and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge for optimal results. AI technology is an interdisciplinary discipline encompassing a wide range of fields, integrating both hardware and software technologies. AI software technologies primarily include natural language processing and machine learning / deep learning.

[0053] Key technologies in speech technology include automatic speech recognition (ASR), text-to-speech (TTS), and voiceprint recognition. Enabling computers to hear, see, speak, and feel is the future direction of human-computer interaction, with speech becoming one of the most promising methods of human-computer interaction.

[0054] From the above, it can be seen that the embodiment of the present application can convert the audio to be classified into audio text. By converting the audio to be classified into audio text, and then classifying the audio text to determine the classification result of the audio to be classified, the fatigue of the classifier in looping the audio to be classified and judging the audio content can be alleviated, and the efficiency of audio classification can be improved.

[0055] The present embodiments are described in detail below. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.

[0056] The embodiment of the present application provides an audio classification method, which can be executed by a terminal or a server, or by both the terminal and the server. The embodiment of the present application takes the audio classification method executed by the terminal as an example for explanation, specifically, the method is executed by an audio classification device integrated in the terminal. Figure 2a As shown, the specific process of the audio classification method can be as follows:

[0057] 201. Display an audio classification page, which includes an audio text converted from the audio to be classified and a classification control for the audio text, wherein the audio text includes a highlighted target text content, which is text content identified from the audio text that matches a preset text content in a text reference database. One classification control corresponds to one classification result.

[0058] Among them, the audio classification page is used to display the audio text of the audio to be classified and the target text content to classify the audio to be classified. The classification control of the audio text in the audio classification page is used to determine the classification result of the audio text.

[0059] The text reference database includes a plurality of preset text contents, wherein the preset text contents may include a plurality of preset types of text contents, such as text contents containing sensitive words, text contents containing joyful words, and so on.

[0060] In one embodiment, in order to improve the classification efficiency of the audio to be classified, the audio to be classified may be converted into audio text and displayed on the audio classification page. The step of "displaying the audio classification page" may include:

[0061] receiving an audio classification request for audio to be classified, and obtaining the audio to be classified based on the audio classification request;

[0062] Perform content recognition on the audio to be classified, and convert the audio to be classified into audio text based on the content recognition result;

[0063] Based on the audio text, display the audio classification page.

[0064] In one example, if Figure 2b As shown in the figure, in addition to displaying the audio text "Son, Aunt Li in quarantine is introducing you to a girl. Come back this weekend to see if she is suitable!" on the audio classification page, you can also display the playback progress bar of the audio to be classified. You can trigger the audio playback control in the playback progress bar to play the audio to be classified. In this example, the target text content can be "Aunt Li", which can be as follows Figure 2b You can also adjust the fill color of the highlight, for example, you can fill it with yellow, red, blue and other colors to distinguish it from the audio text and see it clearly at a glance.

[0065] In one embodiment, in order to further improve the classification efficiency of the audio to be classified, the target text content can be determined from the audio text of the audio to be classified, and then the display form of the target text content is changed to highlighting, and finally displayed on the audio classification page. The step of "displaying the audio classification page" may include:

[0066] Based on the text reference database, the audio text to be classified is identified;

[0067] When it is identified that the audio text to be classified contains target text content containing preset text content in the text reference database, the display form of the target text content is changed to highlighting;

[0068] Displays the audio classification page based on the highlighted results of the target text content.

[0069] In one example, the text reference database includes multiple preset text contents, and each text content in the audio text can be matched with these preset text contents. When there is text content in the audio text that matches the preset text content, the matching text content is highlighted on the audio classification page, and the matching text content is also the target text content.

[0070] 202. In response to a classification operation on a target classification control, determine that the classification control operated by the classification operation is a target classification control, determine a classification result corresponding to the target classification control in the target text content, and determine a classification result of the audio to be classified based on the classification result of the target text content.

[0071] Among them, the target classification control refers to a classification control among multiple classification controls of the audio text, and the target classification control indicates a classification result corresponding to the target text content.

[0072] The classification result of the audio to be classified may be determined based on the classification result of the target text content. For example, when the classification result of the target text content is classification failure, the classification result of the audio to be classified may be determined as classification failure.

[0073] Among them, the response is used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more operations executed can be real-time or have a set delay; unless otherwise specified, there is no restriction on the order of execution of the multiple operations executed.

[0074] In one example, if Figure 2bAs shown, the classification result of the target text content can be determined in response to the classification operation on the classification control 1, and the classification result of the target text content can also be determined in response to the classification operation on the classification control 2. For example, when responding to the classification operation on the classification control 1, the classification result of the target text content can be determined as classification failure, and when responding to the classification operation on the classification control 2, the classification result of the target text content can be determined as classification success.

[0075] In one embodiment, to improve the accuracy of audio classification, after the step of "determining the classification result of the target text content in response to the classification operation on the target classification control, and determining the classification result of the audio to be classified based on the classification result of the target text content", the classification result of the audio to be classified may also be verified. The steps may include:

[0076] When the classification result of the audio to be classified is failed, in response to a trigger operation on the target text content, the target audio corresponding to the target text content in the audio to be classified is played to verify the classification result of the audio to be classified.

[0077] In one example, in order to improve the accuracy of audio classification, when the classification result of the audio to be classified is failed, you can click on the target text to trigger the jump to play the target audio corresponding to the target text content in the audio to be classified, and then you can match the playback result of the target audio with the target text content, and verify the classification result of the audio to be classified based on the matching result. For example, if the playback result of the target audio matches the target text content, it can be verified that the classification result of the audio to be classified is failed. If the playback result of the target audio does not match the target text content, it can be verified that the classification result of the audio to be classified is wrong, and the classification result of the audio to be classified can be changed.

[0078] In one embodiment, the classification page further includes an audio playback control. To improve the accuracy of audio classification, the classification result of the audio to be classified determined based on the classification result of the target text content can be verified. The detailed process of the step "when the classification result of the audio to be classified is a failure, in response to a trigger operation on the target text content, playing the target audio corresponding to the target text content in the audio to be classified to verify the classification result of the audio to be classified" may include:

[0079] When the classification result of the audio to be classified is not passed, in response to a trigger operation on the target text content, determining time information of the target audio corresponding to the target text content in the audio to be classified;

[0080] In response to a trigger operation on the audio playback control, the target audio corresponding to the time information is played to verify the classification result of the audio to be classified.

[0081] In one example, if Figure 2c As shown, in order to improve the accuracy of audio classification, the classification results of the audio to be classified can be verified. For example, the trigger operation of the target text in the audio text can be performed by clicking Figure 2c As shown in the figure, the time information of the target audio corresponding to the text content of "Auntie Li" in the audio to be classified is determined. For example, if the target audio is at 50 seconds of the audio to be classified, the playback progress of the audio to be classified can be jumped to 50 seconds in the playback progress bar. When the trigger operation for the audio playback control is detected, the target audio corresponding to the target text content is played, and then the classification result of the audio to be classified is verified according to the playback result of the target audio and the target text content. The accuracy of audio classification can be improved by playing the target audio corresponding to the target text and then verifying the classification result of the audio to be classified according to the playback result.

[0082] Among them, when it is detected that the target audio corresponding to the target text content has finished playing, the audio classification operation for the target audio is determined, and the classification result of the target audio is obtained, and then the classification type of the target audio is determined. When verifying the classification result of the audio to be classified based on the playback result of the target audio and the target text content, the classification result determined based on the target audio and the classification result of the target text content can be compared. By comparing whether the two classification results are the same, the classification result of the audio to be classified is verified. For example, if it is determined that the classification result of the target audio is type A and the classification result of the target text content is type A, then it can be determined that the classification result of the audio to be classified is correct, and the classification result is also type A, and the verification of the classification result of the audio to be classified is completed.

[0083] In one embodiment, the classification result of the audio to be classified can be verified by matching the playback result of the audio corresponding to the target text content with the target text content to improve the accuracy of audio classification. Specifically, the audio classification method may further include:

[0084] When the playback result of the target audio does not match the target text content, the classification result of the audio to be classified is changed in response to a switching operation on other classification controls, where the other classification controls are controls other than the target classification control in the classification controls.

[0085] Among them, if the playback result of the target audio does not match the target text content, it may be that there are problems such as anomalies when the audio to be classified is converted into audio text, resulting in inaccurate conversion, affecting the classification result of the subsequent target text content, and then affecting the classification result of the audio to be classified. Therefore, the classification result of the audio to be classified can be changed in response to the switching operation of other classification controls, for example, Figure 2dAs shown, when the playback result of the target audio does not match the target text content, the classification result of the audio to be classified is changed in response to the switching operation on the classification control 2.

[0086] Among them, if the playback result of the target audio matches the target text content, the classification result of the audio to be classified does not need to be changed.

[0087] In one embodiment, in order to improve the quality of subsequent audio to be classified and further improve the efficiency of audio classification, descriptive information corresponding to the classification result of the audio to be classified may be sent to the terminal that initiated the audio to be classified. Specifically, the audio classification method may further include:

[0088] Obtain additional information corresponding to the changed classification result of the audio to be classified, the additional information including description information of the changed classification result;

[0089] Based on the description information, a classification result of the audio to be classified is sent to the initiating terminal of the audio to be classified.

[0090] The description information describes the reason for obtaining the classification result of the audio to be classified, such as the audio to be classified contains sensitive information, the audio to be classified contains audio that does not conform to contemporary socialist values, etc.

[0091] In one embodiment, the target text content includes multiple sub-text contents. To improve the accuracy of audio classification, the multiple sub-text contents may be played separately to verify the classification result of the audio to be classified. The audio classification method may further include:

[0092] When the classification result of the audio to be classified is not passed, in response to a trigger operation on the sub-text content, determining that the sub-text content corresponding to the trigger operation is a target sub-text content, and playing a target sub-audio corresponding to the target sub-text content in the audio to be classified;

[0093] When the playback result of the target sub-audio does not match the target sub-text content, in response to the trigger operation for other sub-text contents, the sub-audio corresponding to the other sub-text contents in the audio to be classified is played to verify the classification result of the audio to be classified. The other sub-text contents are the text contents other than the target sub-text content in the multiple sub-text contents.

[0094] Among them, the target text content can be composed of multiple sub-text contents, and the target sub-text content is one of the multiple sub-text contents. For example, there are multiple highlighted text contents in the audio text, and these highlighted text contents can constitute the target text content, and each highlighted text content is the sub-text content mentioned above, and the target sub-text content is one of the multiple highlighted text contents.

[0095] In one example, the target text content includes multiple sub-text contents. When the classification result of the audio to be classified is classification failure, trigger operations can be performed on the multiple sub-text contents one by one, and the corresponding audio can be played to verify the classification result of the audio to be classified.

[0096] The embodiments of the present application can be applied to audio information classification. Traditional audio classification methods usually require information classifiers to play the audio to be classified in a loop and classify the audio files to be classified according to the audio content heard. It is very time-consuming to judge whether the audio information violates the rules, and the efficiency of audio classification is low. The present application can convert the audio to be classified into audio text through speech transcription technology, and then classify the target text content in the audio text in combination with the text reference database. Then, based on the classification result of the target text content, the classification result of the audio to be classified is determined, which can improve the efficiency of audio classification.

[0097] For websites with large amounts of information, audio information needs to be completed by "listening". Through the embodiments of the present application, the specified target content can be quickly obtained while listening to the audio information content, reducing the loss of time invested in audio listening.

[0098] From the above, it can be seen that the embodiment of the present application can convert the audio to be classified into audio text. By converting the audio to be classified into audio text, and then classifying the audio text to determine the classification result of the audio to be classified, the fatigue of the classifier in looping the audio to be classified and judging the audio content can be alleviated, and the efficiency of audio classification can be improved.

[0099] Based on the above introduction, the following examples will be given to further illustrate the audio classification method of this application. Figure 3a , an audio classification method, the specific process can be as follows:

[0100] 301. The terminal receives an audio classification request for audio to be classified, and obtains the audio to be classified based on the audio classification request.

[0101] The audio classification request is initiated by the initiating terminal of the audio to be classified, and the audio classification request can be used to initiate classification of the audio to be classified.

[0102] In one example, if Figure 3bAs shown, after the terminal receives the audio classification request for the audio to be classified, the terminal can open the operation terminal to obtain the audio to be classified, then load the audio to be classified into the operation terminal, and display the audio content. After that, the audio to be classified can be converted into audio text in the background through voice transcription technology, and combined with the text reference database, the target text content in the audio text that matches the preset text content in the text reference database is highlighted. Finally, the audio to be classified, the converted audio text, and the highlighted target text content can be displayed on the audio classification page, and so on.

[0103] 302. The terminal performs content recognition on the audio to be classified, and converts the audio to be classified into audio text based on the content recognition result.

[0104] The terminal may identify the audio content of the audio to be classified, and then convert the identified audio content into text, thereby obtaining the audio text of the audio to be classified.

[0105] In one example, converting the audio to be classified into text can be done by Figure 3b The second step shown is the implementation of speech transcription, that is, the speech is converted into text through the background of the speech transcription technology, and the text is transferred to the operation terminal after transcription. The converted audio text can be displayed on the audio classification page.

[0106] 303. The terminal determines that the audio text contains target text content that matches the preset text content in the text reference database, and highlights the target text content.

[0107] Among them, the text reference database may include multiple preset text contents, and the audio text can be matched with the preset text contents in the text reference database. When there is text content in the audio text that matches the preset text content, the matched text content is used as the target text content, and the target text content is highlighted.

[0108] In one example, if Figure 3b As shown, after semantic transcription, the target text content that matches the preset text content in the text reference database can be highlighted in combination with the text reference database. For example, taking the audio text content "Son, Aunt Li in isolation will introduce you to a girl. Can you come back this weekend to see if she is suitable?" as an example, if the target text content of the audio text is "Aunt Li", "Aunt Li" can be highlighted on the audio classification page, such as by color highlighting.

[0109] 304. The terminal displays an audio classification page based on the audio text and the highlighted target text content. The audio classification page includes a classification control for the audio text.

[0110] Among them, the classification control of the audio text can be used to confirm the classification result of the audio text. There can be multiple classification controls for the audio text. For example, it can include a classification control for confirming that the classification of the audio to be classified has passed, and it can also include a classification control for confirming that the classification of the audio to be classified has failed, and so on.

[0111] In one example, if Figure 3c As shown, the audio classification page can display the playback progress bar and audio playback controls of the audio to be classified. The audio playback controls can be used to play the audio to be classified, and the playback progress bar can prompt the playback progress of the audio to be classified. The playback position of the audio to be classified can be determined by dragging the playback progress bar. The audio classification page can also include the audio text of the audio to be classified, such as Figure 3c The audio classification page can also include classification controls for audio text, for example, Figure 3c As shown, category control 1 and category control 2, and so on.

[0112] 305. The terminal determines a classification result of the target text content in response to a triggering operation on the target classification control, and determines a classification result of the audio to be classified based on the classification result of the target text content.

[0113] Among them, the classification result of the audio to be classified is obtained based on the classification result of the target text content. For example, if the classification result of the target text content is classification failure, then the classification result of the audio to be classified can be determined as classification failure based on the classification failure result of the target text content. If the classification result of the target text content is classification pass, then the classification result of the audio to be classified can be determined as classification pass based on the classification pass result of the target text content.

[0114] In one example, considering the maturity of audio-to-text technology, there may be translation errors. The target audio corresponding to the target text content in the audio to be classified can be obtained, and then the target audio can be played so that the classification result of the audio to be classified can be verified based on the playback result of the target audio and the target text content. For example, the background can record the time node when the target text content appears, and then click the target text content to jump to the time point when the target text content appears in the audio to be classified and play it. When the playback result of the target audio matches the target text content, it can be determined that the classification result of the audio to be classified is correct. If the playback result of the target audio does not match the target text content, it can be determined that the classification result of the audio to be classified is incorrect, and the incorrect classification result can be changed.

[0115] From the above, it can be seen that the embodiment of the present application can convert the audio to be classified into audio text. By converting the audio to be classified into audio text, and then classifying the audio text to determine the classification result of the audio to be classified, the fatigue of the classifier in looping the audio to be classified and judging the audio content can be alleviated, and the efficiency of audio classification can be improved.

[0116] In order to better implement the above method, accordingly, the embodiment of the present application further provides an audio classification device, wherein the audio classification device can be specifically integrated in the server, referring to Figure 4a , the audio classification device may include a page display unit 401 and a result determination unit 402, as follows:

[0117] (1) page display unit 401;

[0118] The page display unit 401 is used to display the audio classification page, which includes the audio text converted from the audio to be classified, and the classification control of the audio text, wherein the audio text includes highlighted target text content, and the target text content is the text content identified from the audio text that matches the preset text content in the text reference database. One classification control corresponds to one classification result.

[0119] In one embodiment, if Figure 4b As shown, the page display unit 401 includes:

[0120] The receiving subunit 4011 is configured to receive an audio classification request for audio to be classified, and obtain the audio to be classified based on the audio classification request;

[0121] The first recognition subunit 4012 is configured to perform content recognition on the audio to be classified and convert the audio to be classified into audio text based on the content recognition result;

[0122] The first page display subunit 4013 is used to display the audio classification page based on the audio text.

[0123] In one embodiment, if Figure 4b As shown, the page display unit 401 includes:

[0124] The second recognition subunit 4014 is used to recognize the audio text to be classified of the audio to be classified based on the text reference database;

[0125] The changing subunit 4015 is configured to change the display mode of the target text content to highlighting when it is identified that the target text content includes the preset text content in the text reference database in the audio text to be classified;

[0126] The second page display subunit 4016 is used to display the audio classification page based on the highlighted result of the target text content.

[0127] (2) result determination unit 402;

[0128] The result determination unit 402 is used to respond to the classification operation on the classification control, determine that the classification control operated by the classification operation is the target classification control, determine the classification result corresponding to the target classification control in the target text content, and determine the classification result of the audio to be classified based on the classification result of the target text content.

[0129] In one embodiment, the audio classification device further includes:

[0130] The first playing unit 403 is used to play the target audio corresponding to the target text content in the audio to be classified when the classification result of the audio to be classified is failed, in response to the trigger operation on the target text content, to verify the classification result of the audio to be classified.

[0131] In one embodiment, if Figure 4c As shown, the first playback unit includes:

[0132] The information determination subunit 4031 is configured to determine the time information of the target audio corresponding to the target text content in the audio to be classified in response to a trigger operation on the target text content when the classification result of the audio to be classified is not passed;

[0133] The playing subunit 4032 is configured to respond to a triggering operation on the audio playing control and play the target audio corresponding to the time information to verify the classification result of the audio to be classified.

[0134] In one embodiment, the audio classification device further includes:

[0135] The changing unit 404 is used to change the classification result of the audio to be classified in response to a switching operation on other classification controls when the playback result of the target audio does not match the target text content. The other classification controls are controls in the classification controls other than the target classification control.

[0136] In one embodiment, the audio classification device further includes:

[0137] An acquiring unit 405 is configured to acquire additional information corresponding to a changed classification result of the audio to be classified, the additional information including description information of the changed classification result;

[0138] The sending unit 406 is configured to send the classification result of the audio to be classified to the initiating terminal of the audio to be classified based on the description information.

[0139] In one embodiment, the audio classification device further includes:

[0140] The second playing unit 407 is configured to, when the classification result of the audio to be classified is a failure, respond to a trigger operation on the sub-text content, determine that the sub-text content corresponding to the trigger operation is a target sub-text content, and play a target sub-audio corresponding to the target sub-text content in the audio to be classified;

[0141] The third playback unit 408 is used to play the sub-audio corresponding to other sub-text content in the audio to be classified in response to a trigger operation for other sub-text content when the playback result of the target sub-audio does not match the target sub-text content, so as to verify the classification result of the audio to be classified. The other sub-text content is the text content other than the target sub-text content in multiple sub-text contents.

[0142] As can be seen from the above, the page display unit 401 of the audio classification device of the embodiment of the present application displays an audio classification page, which includes the audio text after the audio to be classified is converted, and the classification control of the audio text, wherein the audio text includes a highlighted target text content, the target text content is the text content identified from the audio text that matches the preset text content in the text reference database, and one classification control corresponds to one classification result; then, the result determination unit 402 responds to the classification operation on the classification control, determines that the classification control operated by the classification operation is the target classification control, determines the classification result corresponding to the target classification control in the target text content, and determines the classification result of the audio to be classified based on the classification result of the target text content. This scheme can convert the audio to be classified into audio text. By converting the audio to be classified into audio text and then classifying the audio text to determine the classification result of the audio to be classified, the fatigue of the classification personnel who play the audio to be classified in a loop and judge the audio content can be alleviated, and the efficiency of audio classification can be improved.

[0143] In addition, the embodiment of the present application also provides a computer device, which can be a terminal or a server. Figure 5 , which shows a schematic diagram of the structure of the computer device involved in the embodiment of the present application, specifically:

[0144] The computer device may include one or more processing core processors 501, one or more storage media memories 502, a power supply 503, an input unit 504 and other components. Those skilled in the art will understand that Figure 5 The computer device structure shown in the figure does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange components differently.

[0145] Processor 501 is the control center of the computer device. It connects the various components of the entire computer device using various interfaces and lines. By running or executing software programs and / or modules stored in memory 502 and accessing data stored in memory 502, it performs various functions of the computer device and processes data, thereby performing overall testing of the computer device. Optionally, processor 501 may include one or more processing cores; preferably, processor 501 may integrate an application processor and a modem processor, wherein the application processor primarily processes the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 501.

[0146] The memory 502 can be used to store software programs and modules. The processor 501 executes various functional applications and data processing by running the software programs and modules stored in the memory 502. The memory 502 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, at least one application required for a function (such as a sound playback function, an image playback function, etc.); the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 502 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory 502 may also include a memory controller to provide the processor 501 with access to the memory 502.

[0147] The computer device also includes a power supply 503 for supplying power to various components. Preferably, the power supply 503 can be logically connected to the processor 501 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The power supply 503 can also include one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.

[0148] The computer device may further include an input unit 504 , which may be configured to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.

[0149] Although not shown, the computer device may further include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 501 in the computer device will load the executable files corresponding to one or more application processes into the memory 502 according to the following instructions, and the processor 501 will run the application stored in the memory 502 to implement various functions as follows:

[0150] An audio classification page is displayed, which includes an audio text converted from the audio to be classified, and a classification control for the audio text, wherein the audio text includes a highlighted target text content, and the target text content is text content identified from the audio text and matching the preset text content in the text reference database, and one classification control corresponds to one classification result; in response to a classification operation on the classification control, the classification control operated by the classification operation is determined to be a target classification control, the classification result corresponding to the target classification control in the target text content is determined, and based on the classification result of the target text content, the classification result of the audio to be classified is determined.

[0151] From the above, it can be seen that the embodiment of the present application can convert the audio to be classified into audio text. By converting the audio to be classified into audio text, and then classifying the audio text to determine the classification result of the audio to be classified, the fatigue of the classifier in looping the audio to be classified and judging the audio content can be alleviated, and the efficiency of audio classification can be improved.

[0152] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished through instructions, or through instruction-controlled related hardware. The instructions may be stored in a storage medium and loaded and executed by a processor.

[0153] To this end, an embodiment of the present application provides a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute the steps of any of the audio classification methods provided in the embodiments of the present application. For example, the instructions can execute the following steps:

[0154] An audio classification page is displayed, which includes an audio text converted from the audio to be classified, and a classification control for the audio text, wherein the audio text includes a highlighted target text content, and the target text content is text content identified from the audio text and matching the preset text content in the text reference database, and one classification control corresponds to one classification result; in response to a classification operation on the classification control, the classification control operated by the classification operation is determined to be a target classification control, the classification result corresponding to the target classification control in the target text content is determined, and based on the classification result of the target text content, the classification result of the audio to be classified is determined.

[0155] The computer-readable storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0156] Since the instructions stored in the computer-readable storage medium can execute the steps in any audio classification method provided in the embodiments of the present application, the beneficial effects that can be achieved by any audio classification method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.

[0157] According to one aspect of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the audio classification method provided in the above-described summary of the invention and embodiments.

[0158] The above is a detailed introduction to an audio classification method, device, computer equipment and computer-readable storage medium provided in the embodiments of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. An audio classification method, characterized in that: include: Displaying an audio classification page, the audio classification page including an audio text converted from the audio to be classified, a playback progress bar for the audio to be classified, an audio playback control, and a classification control for the audio text, wherein the audio text includes a highlighted target text content, the target text content being text content identified from the audio text that matches preset text content in a text reference database, one classification control corresponding to one audio classification result, the audio text classification control being used to determine the classification result of the audio text, the classification controls including a classification control for confirming that the audio to be classified passes the classification, and a classification control for confirming that the audio to be classified fails the classification; In response to a classification operation on a classification control, determining that the classification control operated by the classification operation is a target classification control, determining a classification result corresponding to the target classification control in the target text content, and determining a classification result for the audio to be classified based on the classification result of the target text content; Before determining the classification result of the audio to be classified based on the classification result of the target text content, the method further includes: When the classification result of the audio to be classified is failure, in response to a triggering operation on the target text content, determining time information of the target audio corresponding to the target text content in the audio to be classified; In response to a triggering operation on the audio playback control, the target audio corresponding to the time information is played to verify a classification result of the audio to be classified.

2. The method according to claim 1, characterized in that The method further comprises: When the playback result of the target audio does not match the target text content, the classification result of the audio to be classified is changed in response to a switching operation on other classification controls, where the other classification controls are controls in the classification controls other than the target classification control.

3. The method according to claim 2, characterized in that The method further comprises: Acquire additional information corresponding to a changed classification result of the audio to be classified, the additional information including description information of the changed classification result; Based on the description information, a classification result of the audio to be classified is sent to an initiating terminal of the audio to be classified.

4. The method according to claim 1, wherein The audio classification page includes: receiving an audio classification request for audio to be classified, and obtaining the audio to be classified based on the audio classification request; Performing content recognition on the audio to be classified, and converting the audio to be classified into audio text based on the content recognition result; Based on the audio text, an audio classification page is displayed.

5. The method according to claim 1, wherein The audio classification page includes: Based on the text reference database, the audio text to be classified is identified; When it is identified that the audio text to be classified contains target text content containing preset text content in the text reference database, changing the display form of the target text content to highlighting; Based on the highlighting result of the target text content, the audio classification page is displayed.

6. The method according to claim 1, characterized in that The target text content includes a plurality of sub-text contents, and the method further includes: When the classification result of the audio to be classified is not passed, in response to a trigger operation on a sub-text content, determining that the sub-text content corresponding to the trigger operation is a target sub-text content, and playing a target sub-audio corresponding to the target sub-text content in the audio to be classified; When the playback result of the target sub-audio does not match the target sub-text content, in response to a trigger operation for other sub-text contents, the sub-audio corresponding to the other sub-text content in the audio to be classified is played to verify the classification result of the audio to be classified, and the other sub-text content is the text content in the multiple sub-text contents except the target sub-text content.

7. An audio classification device, characterized in that: include: A page display unit, configured to display an audio classification page, the audio classification page including an audio text converted from the audio to be classified, a playback progress bar for the audio to be classified, an audio playback control, and a classification control for the audio text, wherein the audio text includes a highlighted target text content, the target text content being text content identified from the audio text that matches preset text content in a text reference database, one classification control corresponding to one audio classification result, the classification control for the audio text being used to determine the classification result of the audio text, the classification controls including a classification control for confirming that the audio to be classified has passed the classification, and a classification control for confirming that the audio to be classified has failed the classification; A result determination unit is configured to, in response to a classification operation on a classification control, determine that the classification control operated by the classification operation is a target classification control, determine a classification result corresponding to the target classification control in the target text content, and determine a classification result of the audio to be classified based on the classification result of the target text content; The device further includes a first playback unit, which includes: an information determination subunit, configured to, when the classification result of the audio to be classified is a failure, determine, in response to a trigger operation on the target text content, time information of the target audio corresponding to the target text content in the audio to be classified; The playing subunit is used to respond to the trigger operation of the audio playing control and play the target audio corresponding to the time information to verify the classification result of the audio to be classified.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the audio classification method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Teaching record data correcting device

    CN107220228A

  • Voice quality inspection display method and device and electronic equipment

    CN110750229A

  • Manuscript display control method and device, electronic equipment and storage medium

    CN111970257A