Terminal device and audio data analysis method
By performing sampling rate alignment and visual waveform comparison on the terminal device, the problem of inaccurate results in existing audio data analysis methods is solved, and accurate analysis in specific scenarios is achieved.
Patent Information
- Application Number
- CN202310532427.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-11
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2043-05-11
Smart Images

Figure CN118942450B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminal equipment technology, and in particular to a terminal equipment and a method for analyzing audio data. Background Technology
[0002] To understand the characteristics of a particular type of data, model analysis can be used. For example, this type of data can be input into a classification model, and the characteristics of this type of data can be obtained through the analysis of the classification model.
[0003] Taking audio data analysis as an example, audio data can be input into an audio classification model. Through the reasoning and prediction of the audio classification model, a prediction result is generated, which is the data feature of the audio data.
[0004] However, when analyzing audio data using audio classification models, only the overall performance of the audio data can be analyzed; that is, only the global performance of the audio data can be analyzed, and it cannot perform targeted analysis on data in specific scenarios. For example, taking voice wake-up in a voice wake-up model as an example, after inputting a wake-up word, the voice wake-up model can only analyze the overall performance of the wake-up word, but cannot analyze the performance of data in specific scenarios. For example, the data performance of a user at fast and slow speaking speeds, or the data performance of a specific timbre, etc., the voice wake-up model can only output abstract effects and cannot intuitively analyze the data performance. Therefore, there is a problem of inaccurate results when analyzing data using model analysis methods. Summary of the Invention
[0005] Some embodiments of this application provide a terminal device and an audio data analysis method to solve the problem of inaccurate results when performing data analysis using model analysis methods.
[0006] In a first aspect, some embodiments of this application provide a terminal device, including:
[0007] A memory, wherein program instructions are stored;
[0008] The processor, by executing the program instructions, is configured to:
[0009] Input the test audio file into the preset model;
[0010] Obtain the confidence score result obtained by the preset model performing inference on the test audio file, and store the confidence score result as a confidence score list;
[0011] The sampling rate of the test audio file is calculated according to the inference step size of the preset model;
[0012] The confidence list is saved as a target audio file based on the sampling rate.
[0013] The audio wave is extracted from the test audio file, and the target audio file is imported into the first audio track to obtain the confidence waveform of the target audio file;
[0014] The audio wave and the confidence waveform are output in a visual form;
[0015] By comparing the audio sound wave with the confidence waveform, waveform difference data is generated, and analysis results of the test audio file are generated based on the waveform difference data.
[0016] As can be seen from the above technical solutions, some embodiments of this application provide a terminal device. The terminal device aligns the test audio file with the analysis results of the preset model in the time domain by means of the sampling rate, saves the confidence list as the target audio file, and outputs the analysis results in a visual form. This can solve the problem that the results are not intuitive and inaccurate when performing data analysis by model analysis methods.
[0017] Secondly, some embodiments of this application provide an audio data analysis method, which can be applied to the terminal device of the first aspect. The audio data analysis method includes:
[0018] Input the test audio file into the preset model;
[0019] Obtain the confidence score result obtained by the preset model performing inference on the test audio file, and store the confidence score result as a confidence score list;
[0020] The sampling rate of the test audio file is calculated according to the inference step size of the preset model;
[0021] The confidence list is saved as a target audio file based on the sampling rate.
[0022] The audio wave is extracted from the test audio file, and the target audio file is imported into the first audio track to obtain the confidence waveform of the target audio file;
[0023] The audio wave and the confidence waveform are output in a visual form;
[0024] By comparing the audio sound wave with the confidence waveform, waveform difference data is generated, and analysis results of the test audio file are generated based on the waveform difference data.
[0025] As can be seen from the above technical solutions, some embodiments of this application provide an audio data analysis method. The method aligns the analysis results of the test audio file and the preset model in the time domain by sampling rate, saves the confidence list as the target audio file, and outputs the analysis results in a visual form. This can solve the problem of inaccurate results when performing data analysis through model analysis methods. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in some embodiments of this application or in the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This is a schematic diagram illustrating the operational scenarios between a terminal device and a control device provided in some embodiments of this application;
[0028] Figure 2 Hardware configuration block diagrams of terminal devices provided in some embodiments of this application;
[0029] Figure 3 Hardware configuration block diagrams of control devices provided in some embodiments of this application;
[0030] Figure 4 This is a schematic diagram of the software configuration in a terminal device provided in some embodiments of this application;
[0031] Figure 5 A schematic diagram illustrating the confidence level output of the audio classification model provided in some embodiments of this application;
[0032] Figure 6 This application provides a schematic flowchart of a terminal device performing an audio data analysis method according to some embodiments;
[0033] Figure 7 A schematic diagram illustrating the process of a terminal device obtaining confidence results according to some embodiments of this application;
[0034] Figure 8 A schematic diagram illustrating the process of a terminal device storing confidence results as a confidence list, provided in some embodiments of this application;
[0035] Figure 9 A schematic diagram illustrating the process of a terminal device, provided in some embodiments of this application, calculating the sampling rate of a test audio file according to the inference step size of a preset model;
[0036] Figure 10 A schematic diagram illustrating the process of a terminal device saving a confidence list as a target audio file, provided in some embodiments of this application;
[0037] Figure 11 A schematic diagram illustrating the process of a terminal device extracting audio waves from a test audio file, provided in some embodiments of this application;
[0038] Figure 12 This is a schematic diagram illustrating the effect of a terminal device importing a target audio file into a first audio track, provided in some embodiments of this application.
[0039] Figure 13 A schematic diagram illustrating the process of a terminal device provided in some embodiments of this application outputting audio sound waves and confidence waveforms in a visual form;
[0040] Figure 14 A schematic diagram illustrating the effect of visualizing output audio sound waves and confidence waveforms in the same display channel in some embodiments of this application;
[0041] Figure 15 A schematic diagram illustrating the process of generating waveform difference data by comparing audio sound waves and confidence waveforms in some embodiments of this application;
[0042] Figure 16a A schematic diagram illustrating the visualization effect of waveform difference data provided in some embodiments of this application;
[0043] Figure 16b This is a schematic diagram illustrating another visualization effect of waveform difference data provided in some embodiments of this application;
[0044] Figure 17 A schematic diagram illustrating the process of marking the total display position in the display channel for a terminal device provided in some embodiments of this application;
[0045] Figure 18 A schematic diagram illustrating the output effect of audio sound waves and confidence waveforms at a first resolution provided in some embodiments of this application;
[0046] Figure 19 A schematic diagram illustrating the output effect of audio sound waves and confidence waveforms at a second resolution, provided for some embodiments of this application;
[0047] Figure 20 This is a schematic flowchart illustrating an audio data analysis method provided in some embodiments of this application. Detailed Implementation
[0048] To make the objectives, technical solutions, and advantages of some embodiments of this application clearer, the technical solutions of some embodiments of this application will be clearly and completely described below with reference to specific embodiments and corresponding drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0049] It should be noted that the brief descriptions of terms in some embodiments of this application are only for the convenience of understanding the implementation methods described below, and are not intended to limit the implementation methods of some embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.
[0050] In some embodiments of this application, the terms "first," "second," "third," etc., in the specification, claims, and accompanying drawings are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms can be used interchangeably where appropriate.
[0051] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.
[0052] The term "module" refers to any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.
[0053] Figure 1 This is a schematic diagram illustrating an operational scenario between a terminal device and a control device provided in some embodiments of this application. For example... Figure 1 As shown, a user can operate a terminal device 200 via a mobile terminal 300 and a control device 100.
[0054] In some embodiments, the control device 100 may be a remote control. Communication between the remote control and the terminal device includes infrared protocol communication, Bluetooth protocol communication, and other short-range communication methods, controlling the terminal device 200 wirelessly or via other wired means. Users can input user commands through buttons on the remote control, voice input, control panel input, etc., to control the terminal device 200.
[0055] In some embodiments, the mobile terminal 300 can install software applications with the terminal device 200 to establish a connection and communication via a network communication protocol, thereby achieving one-to-one control operations and data communication. Alternatively, audio and video content displayed on the mobile terminal 300 can be transmitted to the terminal device 200 to achieve synchronous display.
[0056] like Figure 1 The document also shows that terminal device 200 can communicate with server 400 via various communication methods. Terminal device 200 can communicate via local area network (LAN), wireless local area network (WLAN), and other networks.
[0057] In addition to providing broadcast television reception functions, terminal device 200 can also be equipped with intelligent network television functions that provide computer support, including but not limited to network television, smart television, Internet Protocol television (IPTV), etc.
[0058] Figure 2 This is a hardware configuration block diagram of a terminal device provided in some embodiments of this application.
[0059] In some embodiments, the terminal device 200 includes at least one of a tuner 210, a communicator 220, a detector 230, an external device interface 240, a controller 250, a display 260, an audio output interface 270, a memory, a power supply, and a user interface.
[0060] In some embodiments, the detector 230 is used to collect signals from the external environment or to interact with the outside world.
[0061] In some embodiments, the display 260 includes a display screen component for presenting an image, a driving component for driving image display, a component for receiving image signals output from a controller, and a user control UI interface, etc.
[0062] In some embodiments, the communicator 220 is a component for communicating with external devices or servers according to various communication protocol types.
[0063] In some embodiments, the controller 250 controls the operation of the terminal device and responds to user operations through various software control programs stored in the memory. The controller 250 controls the overall operation of the terminal device 200.
[0064] In some embodiments, a user can input user commands through a graphical user interface (GUI) displayed on a display 260, and the user input interface receives user input commands through the graphical user interface (GUI).
[0065] In some embodiments, user interface 280 is an interface that can be used to receive control input.
[0066] Figure 3 This is a hardware configuration block diagram of a control device provided in some embodiments of this application. For example... Figure 3 As shown, the control device 100 includes a controller 110, a communication interface 130, a user input / output interface, a memory, and a power supply.
[0067] The control device 100 is configured to control the terminal device 200, and can receive user input operation commands and convert the operation commands into commands that the terminal device 200 can recognize and respond to, thus acting as an intermediary for interaction between the user and the terminal device 200.
[0068] In some embodiments, the control device 100 may be an intelligent device. For example, the control device 100 may be equipped with various applications of the control terminal device 200 according to user needs.
[0069] In some embodiments, such as Figure 1 As shown, the mobile terminal 300 or other smart electronic devices can perform similar functions to the control device 100 after the application of the control terminal device 200 is installed.
[0070] The controller 110 includes a processor 112, RAM 113, ROM 114, a communication interface 130, and a communication bus. The controller 110 is used to control the operation of the control device 100, as well as the communication and cooperation between internal components and the external and internal data processing functions.
[0071] Under the control of the controller 110, the communication interface 130 enables communication of control signals and data signals with the terminal device 200. The communication interface 130 may include at least one of other near-field communication modules such as WiFi chip 131, Bluetooth module 132, and NFC module 133.
[0072] User input / output interface 140, wherein the input interface includes at least one of other input interfaces such as microphone 141, touchpad 142, sensor 143, button 144, etc.
[0073] In some embodiments, the control device 100 includes at least one of a communication interface 130 and an input / output interface 140. The control device 100 is configured with the communication interface 130, such as a WiFi, Bluetooth, or NFC module, which can encode user input commands via WiFi, Bluetooth, or NFC protocols and send them to the terminal device 200.
[0074] The memory 190 is used to store various operating programs, data, and applications for driving and controlling the control device 100 under the control of the controller. The memory 190 can also store various control signal instructions input by the user.
[0075] The power supply 180 is used to provide operating power support for the various components of the control device 100 under the control of the controller.
[0076] Figure 4The diagram illustrates the software configuration in a terminal device according to some embodiments of this application. In some embodiments, the system is divided into four layers, from top to bottom: the Applications layer (hereinafter referred to as the "Application Layer"), the Application Framework layer (hereinafter referred to as the "Framework Layer"), the Android runtime and the System Library layer (hereinafter referred to as the "System Runtime Library Layer"), and the kernel layer.
[0077] In some embodiments, at least one application runs in the application layer. These applications may be Windows programs that come with the operating system, system settings programs, clock programs, camera applications, etc., or they may be applications developed by third-party developers.
[0078] The framework layer provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. The application framework layer includes predefined functions. It acts as a central processing unit, determining the actions taken by applications within the application layer.
[0079] like Figure 4 As shown, the application framework layer in this embodiment includes managers, content providers, and a view system.
[0080] In some embodiments, the Activity Manager is used to: manage the lifecycle of individual applications and the usual navigation back functionality.
[0081] In some embodiments, the window manager is used to manage all window programs.
[0082] In some embodiments, the system runtime library layer provides support for the upper layer, namely the framework layer. When the framework layer is accessed, the Android operating system runs the C / C++ libraries contained in the system runtime library layer to implement the functions that the framework layer needs to perform.
[0083] In some embodiments, the kernel layer is a layer between hardware and software. For example... Figure 4 As shown, the kernel layer includes at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, touch sensor, pressure sensor, etc.).
[0084] In some embodiments, the kernel layer also includes a power driver module for power management.
[0085] In some embodiments, Figure 4 The software programs and / or modules corresponding to the software architecture in the document are stored in [the relevant database]. Figure 2 or Figure 3 In the first or second memory shown.
[0086] Based on the aforementioned terminal device 200, media data can be played and identified through the terminal device 200. Different media data categories correspond to different data characteristics. To understand the characteristics of a certain type of data, model analysis methods can be used. For example, the type of data to be understood can be input into a classification model, and through the analysis, reasoning, and prediction of the classification model, the characteristics of that type of data can be obtained.
[0087] For example, taking audio data from media data as an example, in order to understand the characteristics of audio data, the audio data can be input into an audio classification model. Audio data may contain different timbres, different speech rates, etc. Through the reasoning and prediction of the audio classification model, a prediction result for the audio data can be generated, and this prediction result is the data feature of the audio data.
[0088] Figure 5 This is a schematic diagram illustrating the confidence level output of the audio classification model provided in some embodiments of this application, such as... Figure 5 As shown, in some embodiments, when analyzing audio data using an audio classification model, only the overall performance of the audio data can be analyzed, that is, only the global performance of the audio data can be analyzed, and targeted analysis of specific scene data is not possible. During the iterative optimization of test data using an audio classification model, it is necessary to perform analysis of audio features and related sound scenes based on the performance of the audio classification model on the test data. However, as... Figure 5 As shown, the audio classification model does not perform well on test data and is not very sensitive to features. It can only obtain time information and the overall situation of confidence. If we want to analyze the performance of the audio classification model on different features and sound scenes, we need to find the pattern of confidence changes.
[0089] For example, taking voice wake-up in a voice wake-up model as an example, after inputting a wake-up word, the voice wake-up model can only analyze the overall performance of the wake-up word, and cannot analyze the performance of data in specific scenarios. For example, the voice wake-up model can only output abstract effects on the user's data performance at fast and slow speaking speeds, or the data performance of a specific timbre. For specific scenarios such as not being woken up or being mistakenly woken up, it cannot intuitively analyze the performance of the data. Therefore, there is a problem of inaccurate results when analyzing data using model analysis methods.
[0090] To address the problem of inaccurate results when performing data analysis using model analysis methods, some embodiments of this application provide a terminal device 200. The terminal device 200 may include a memory 190 and a processor 112. The memory 190 stores program instructions, and the processor 112 can execute the program instructions stored in the memory 190. To facilitate understanding of the technical solutions in some embodiments of this application, the steps are described in detail below with reference to specific embodiments and accompanying drawings. Figure 6 This application provides schematic diagrams illustrating the audio data analysis method performed by a terminal device in some embodiments. Figure 6 As shown, when performing audio data analysis, the terminal device 200 may include the following steps S1-S7, the specific contents of which are as follows:
[0091] Step S1: Terminal device 200 inputs a test audio file to the preset model.
[0092] In some embodiments, the preset model can be an audio classification model. To understand the audio data characteristics of the test audio file, the test audio file can be input into the audio classification model. After input, the audio classification model can perform inference and prediction on the test audio file to generate corresponding inference results. After step S1 is completed, step S2 can be executed.
[0093] Step S2: The terminal device 200 obtains the confidence results obtained by performing inference on the test audio file using a preset model, and stores the confidence results as a confidence list.
[0094] After the test audio file is input into the preset model, the preset model can perform inference on the test audio file to obtain the confidence result. Figure 7 This is a schematic diagram illustrating the process of a terminal device obtaining confidence results according to some embodiments of this application, such as... Figure 7 As shown, in some embodiments, the terminal device 200 can obtain the confidence score result obtained by performing inference on the test audio file using a preset model in the following manner. The terminal device 200 first generates a request instruction for obtaining the confidence score result. In response to the request instruction, it runs the preset model and returns the confidence score result of the test audio file according to the request instruction. The terminal device 200 then obtains the output format of the confidence score result based on the returned confidence score result and outputs the confidence score result according to the output format. The output format of the confidence score result can be set according to actual needs, and this application does not specifically limit it.
[0095] To facilitate the intuitive display of the confidence results of the test audio file later, in some embodiments, the terminal device 200 can store the confidence results as a confidence list after obtaining the confidence results of the test audio file. Figure 8This application provides a schematic diagram illustrating the process of a terminal device storing confidence results as a confidence list in some embodiments, as shown below. Figure 8 As shown, when storing the confidence results as a confidence list, the terminal device 200 can first iterate through the confidence results, then create a confidence list for storing the confidence results based on the output format of the confidence results, and then add the confidence results to the confidence list.
[0096] For example, when iterating through the confidence scores, the attribute information contained in the confidence scores can be obtained. For instance, the confidence scores may contain multiple fields. Then, a confidence score list for storing the confidence scores is created by combining the multiple fields and the output format of the confidence scores. After the confidence score list is created, the confidence scores are added to the confidence score list. After step S2 is completed, step S3 can be executed.
[0097] Step S3: The terminal device 200 calculates the sampling rate of the test audio file according to the inference step size of the preset model.
[0098] In order to ensure that the analysis results of the test audio file and the preset model are aligned in the time domain, in some embodiments, the terminal device 200 can calculate the sampling rate of the test audio file according to the inference step size of the preset model, so as to achieve the alignment of the analysis results of the preset model and the test audio file in the time domain. Figure 9 This application provides a schematic diagram illustrating the process of a terminal device calculating the sampling rate of a test audio file according to the inference step size of a preset model, as shown in some embodiments. Figure 9 As shown, when calculating the sampling rate of a test audio file, the terminal device 200 first obtains the unit time used for sampling the test audio file and the inference step size of the preset model. Then, based on the ratio of the unit time to the inference step size, it calculates the sampling rate of the test audio file. It should be noted that there is no fixed order between obtaining the unit time and obtaining the inference step size of the preset model; that is, the unit time can be obtained first and then the inference step size, or the inference step size can be obtained first and then the unit time, or both can be obtained simultaneously. This application does not limit this.
[0099] For example, with a unit time of 1 second and a step size of 100 milliseconds, after obtaining a unit time of 1 second (1000 milliseconds) and a step size of 100 milliseconds, the sampling rate of the test audio file can be calculated using the ratio of the unit time to the inference step size, i.e., the ratio of 1000 milliseconds to 100 milliseconds. In other words:
[0100] The sampling rate of the test audio file = unit time / inference step = 1000 milliseconds / 100 milliseconds = 10
[0101] That is, the sampling rate needs to be set to 10, sampling 10 points per second. By setting the sampling rate of the test audio file, it can be ensured that the analysis results of the test audio file and the preset model correspond to each other, that is, that the analysis results of the test audio file and the preset model are aligned in the time domain. After the sampling rate calculation in step S3 is completed, step S4 can be executed.
[0102] Step S4: The terminal device 200 saves the confidence list as the target audio file based on the sampling rate of the test audio file.
[0103] To facilitate a clear and intuitive display of the confidence results of the test audio file, in some embodiments, after the sampling rate of the test audio file is calculated, the terminal device 200 can save the confidence list as a target audio file based on the sampling rate of the test audio file. Figure 10 This application provides a schematic diagram illustrating the process of a terminal device saving a confidence list as a target audio file, as shown in some embodiments. Figure 10 As shown, when the terminal device 200 saves the confidence list as a target audio file, it can first traverse the confidence list to obtain the initial format of the confidence list, then obtain the target audio format of the target audio file, and finally convert the initial format into the target audio format to save the confidence list as the target audio file.
[0104] For example, the target audio file is in the standard digital audio file format WAV. WAV format is a lossless music format that more easily preserves the quality of the original audio. After obtaining the initial format of the confidence list, the terminal device 200 can save the confidence list as a WAV file to ensure that the audio in the test audio file is not damaged. In actual operation, the confidence list can be saved as the target audio file using an audio conversion tool. For example, the confidence list can be input into an audio conversion tool, which will then perform the conversion operation and output the corresponding conversion result. Other methods can also be used to save the confidence list as the target audio file; the specific method is not limited in this application.
[0105] It should be noted that when the terminal device 200 saves the confidence list as the target audio file, it does so only after calculating the sampling rate of the test audio file according to the inference step size of the preset model. In other words, the conversion is performed only after the test audio file and the analysis results of the preset model are aligned in the time domain. This ensures the accuracy of the analysis results of the test audio file and solves the problem of inaccuracies that can occur when analyzing data using model-based analysis methods. After step S4 is completed, step S5 can be executed.
[0106] Step S5: The terminal device 200 extracts the audio wave from the test audio file and imports the target audio file into the first audio track to obtain the confidence waveform of the target audio file.
[0107] In order to obtain the audio features in the test audio file, the terminal device 200 can extract the audio sound waves from the test audio file. Figure 11 This application provides schematic diagrams illustrating the process of a terminal device extracting audio waves from a test audio file, as shown in some embodiments. Figure 11 As shown, when the terminal device 200 extracts audio sound waves from the test audio file, it can first parse the test audio file to obtain the audio amplitude information of the test audio file, then traverse the audio amplitude information to obtain the vibration range of the audio amplitude information, and then generate the audio sound waves of the test audio file based on the vibration range.
[0108] For example, when extracting audio sound waves, a test audio file can be input into audio analysis software, which then parses the file. The analysis generates audio amplitude information from the test audio file. It's understood that audio amplitude information has a certain vibration range; a larger vibration range indicates greater audio fluctuation in the test audio file, while a smaller range indicates less fluctuation. Based on this vibration range, the audio sound waves from the test audio file can be generated.
[0109] In order to obtain the audio features in the target audio file, the terminal device 200 can import the target audio file into the first audio track to obtain the confidence waveform of the target audio file. Figure 12 This application provides schematic diagrams illustrating the effect of a terminal device importing a target audio file into a first audio track, as shown in some embodiments. Figure 12 As shown, since the target audio file is in a common audio format, its confidence waveform can be visually presented through audio analysis applications. This allows users to directly observe the confidence waveform of the target audio file, eliminating the need to analyze and summarize patterns from numerous confidence results individually. This solves the problem of unintuitive results when performing data analysis using model-based methods and also improves the processing efficiency of the terminal device 200. After step S5 is completed, step S6 can be executed.
[0110] Step S6: The terminal device 200 outputs the audio sound wave and confidence waveform in a visual form.
[0111] After the terminal device 200 extracts the audio wave from the test audio file and imports the target audio file into the first track to obtain the confidence waveform of the target audio file, the terminal device 200 can output the audio wave and confidence waveform in a visual form. Figure 13This application provides a schematic diagram illustrating the process of a terminal device outputting audio sound waves and confidence waveforms in a visual form, as shown in some embodiments. Figure 13 As shown, the terminal device 200 first obtains the display channel where the confidence waveform is located, i.e., the display channel of the first audio track. Then, it creates a second audio track in this display channel and imports the audio wave into the second audio track. In this way, the first audio track where the confidence waveform of the target audio file is located and the second audio track where the audio wave of the test audio file is located will be displayed in the same display channel. This allows for the simultaneous display of audio waves and confidence waveforms in the same display channel but on different audio tracks, i.e., the simultaneous output of audio waves and confidence waveforms in a visual form.
[0112] For example, Figure 14 This application provides schematic diagrams illustrating the effect of visualizing output audio sound waves and confidence waveforms in the same display channel in some embodiments, such as... Figure 14 As shown, in the same display channel, the confidence waveform is displayed on the first audio track, and the audio wave is displayed on the second audio track. This allows for a direct understanding of the characteristics of the audio data through the confidence waveform of the target audio file and the audio wave in the test audio file. After step S6 is completed, step S7 can be executed.
[0113] Step S7: The terminal device 200 compares the audio sound wave with the confidence waveform, generates waveform difference data, and generates analysis results of the test audio file based on the waveform difference data.
[0114] In order to obtain the difference between the audio sound wave and the confidence waveform to identify problematic audio data, in some embodiments, after obtaining the confidence waveform of the target audio file and the audio sound wave in the test audio file from step S6, the terminal device 200 compares the audio sound wave and the confidence waveform to generate waveform difference data, so as to generate the analysis result of the test audio file based on the waveform difference data, that is, to determine the problematic audio data based on the waveform difference data.
[0115] Figure 15 This is a schematic diagram illustrating the process of a terminal device generating waveform difference data by comparing audio sound waves and confidence waveforms, as provided in some embodiments of this application. Figure 15 As shown, when generating waveform difference data, the terminal device 200 can first traverse the audio amplitude information of the test audio file, then traverse the confidence waveform to obtain the waveform amplitude information of the confidence waveform, and finally compare the audio amplitude information and the waveform amplitude information to obtain the waveform difference data.
[0116] For example, when the terminal device 200 traverses the audio amplitude information of the test audio file, it can obtain the audio amplitude range of the test audio file. When traversing the confidence waveform, it can obtain the vibration range of the confidence waveform. Then, it can compare the audio amplitude range of the test audio file with the vibration range of the confidence waveform. Where the range or vibration amplitude of the two do not match, it can be determined that it is waveform difference data.
[0117] Figure 16a This is a schematic diagram illustrating the visualization effect of waveform difference data provided in some embodiments of this application. Figure 16b This is a schematic diagram illustrating another visualization effect of waveform difference data provided in some embodiments of this application, such as... Figure 16a , Figure 16b As shown, by comparing the confidence waveform of the target audio file in the first audio track and the audio wave of the test audio file in the second audio track in the same display channel, it can be clearly seen that there is a mismatch between the confidence waveform and the audio wave amplitude, i.e., waveform difference data.
[0118] In order to make it intuitive for users to see the location of waveform difference data, in some embodiments, the terminal device 200 can mark the range of waveform difference data. Figure 17 This is a schematic diagram illustrating the process of marking the total display position in the display channel of a terminal device provided in some embodiments of this application, such as... Figure 17 As shown, the terminal device 200 can first locate the waveform difference data, that is, obtain the first display position where the audio amplitude information of the test audio file is located and the second display position where the waveform amplitude information of the confidence waveform is located. Since the two are in the same display channel, the first display position and the second display position can be merged to generate the total display position of the waveform difference data, and the total display position is marked in the display channel, thus realizing the marking of the total display position of the waveform difference data, as shown. Figure 16a , Figure 16b The area within the box is shown in the image. By comparing the audio sound wave and the confidence waveform to generate waveform difference data, it's possible to pinpoint the specific part of the audio data in the test audio file where the problem exists, rather than simply reflecting the overall characteristics of the audio. This improves the accuracy of test audio file analysis and solves the problem of inaccurate results when performing data analysis using model-based methods.
[0119] In order to more accurately output the location of waveform difference data, in some embodiments, the terminal device 200 may also provide output of waveform difference data at different resolutions. Figure 18 This is a schematic diagram illustrating the output effect of audio sound waves and confidence waveforms at a first resolution, provided in some embodiments of this application. Figure 19This is a schematic diagram illustrating the output effect of audio sound waves and confidence waveforms at a second resolution provided in some embodiments of this application. The resolutions of the audio sound waves and confidence waveforms are set differently in the two cases. Figure 18 At the resolution shown, the output of the audio sound wave and confidence waveform is relatively smooth. Figure 19 At the resolution shown, the output of the audio wave and confidence waveform is more refined. This allows users to set the resolution of the audio wave and confidence waveform according to their actual needs, enabling more accurate determination of waveform differences and further improving the accuracy of test audio file analysis.
[0120] As can be seen from the above technical solutions, the above embodiments provide a terminal device. The terminal device inputs a test audio file into a preset model, obtains the confidence result obtained by the preset model performing inference on the test audio file, and stores the confidence result as a confidence list; calculates the sampling rate of the test audio file according to the inference step size of the preset model, and saves the confidence list as a target audio file based on the sampling rate; extracts audio waves from the test audio file, and imports the target audio file into a first audio track to obtain the confidence waveform of the target audio file; outputs the audio waves and confidence waveforms in a visual form, compares the audio waves and confidence waveforms to generate waveform difference data, and generates the analysis result of the test audio file based on the waveform difference data. The terminal device aligns the test audio file with the analysis result of the preset model in the time domain by using the sampling rate, saves the confidence list as the target audio file, and outputs the analysis result in a visual form, which can solve the problem of inaccurate results when performing data analysis using model analysis methods.
[0121] This application also provides a method for analyzing audio data in some embodiments, which can be applied to the terminal device 200 in the above embodiments. Figure 20 This is a schematic flowchart of an audio data analysis method provided in some embodiments of this application, such as... Figure 20 As shown, in some embodiments, an audio data analysis method may include the following steps S1-S7, the details of which are as follows:
[0122] Step S1: Terminal device 200 inputs a test audio file to the preset model.
[0123] In some embodiments, the preset model can be an audio classification model. To understand the audio data characteristics of the test audio file, the test audio file can be input into the audio classification model. After input, the audio classification model can perform inference and prediction on the test audio file to generate corresponding inference results. After step S1 is completed, step S2 can be executed.
[0124] Step S2: The terminal device 200 obtains the confidence results obtained by performing inference on the test audio file using a preset model, and stores the confidence results as a confidence list.
[0125] After a test audio file is input into a preset model, the preset model can perform inference on the test audio file to obtain a confidence score result. In some embodiments, the terminal device 200 can obtain the confidence score result obtained by the preset model performing inference on the test audio file in the following manner. For example, the terminal device 200 can first generate a request instruction for obtaining the confidence score result. In response to the request instruction, the preset model receives the request instruction and returns the confidence score result of the test audio file according to the request instruction. The terminal device 200 receives the confidence score result returned by the preset model according to the request instruction, obtains the output format of the confidence score result, and then outputs the confidence score result according to the output format. The output format of the confidence score result can be set according to actual needs, and this application does not specifically limit it.
[0126] To facilitate a clear and intuitive display of the confidence results, in some embodiments, after obtaining the confidence results of the test audio file, the terminal device 200 can store the confidence results as a confidence list. When storing the confidence results as a confidence list, the terminal device 200 first iterates through the confidence results, then creates a confidence list for storing the confidence results based on the output format of the confidence results, and finally adds the confidence results to the confidence list. After step S2 is completed, step S3 can be executed.
[0127] Step S3: The terminal device 200 calculates the sampling rate of the test audio file according to the inference step size of the preset model.
[0128] To ensure temporal alignment between the analysis results of the test audio file and the preset model, in some embodiments, the terminal device 200 can calculate the sampling rate of the test audio file according to the inference step size of the preset model, thereby achieving temporal alignment between the analysis results of the preset model and the test audio file. When calculating the sampling rate of the test audio file, the terminal device 200 first obtains the unit time used for sampling the test audio file and the inference step size of the preset model. Then, based on the ratio of the unit time to the inference step size, it calculates the sampling rate of the test audio file. It should be noted that there is no fixed order between obtaining the unit time and obtaining the inference step size of the preset model; that is, the unit time can be obtained first and then the inference step size, or the inference step size can be obtained first and then the unit time, or both can be obtained simultaneously. This application does not limit this. After the sampling rate calculation in step S3 is completed, step S4 can be executed.
[0129] Step S4: The terminal device 200 saves the confidence list as the target audio file based on the sampling rate of the test audio file.
[0130] To facilitate a clear and intuitive display of the confidence scores of the test audio files, in some embodiments, after the sampling rate of the test audio files is calculated, the terminal device 200 can save the confidence score list as a target audio file based on the sampling rate of the test audio files. When saving the confidence score list as a target audio file, the terminal device 200 first iterates through the confidence score list to obtain its initial format, then obtains the target audio format of the target audio file, and finally converts the initial format to the target audio format to save the confidence score list as the target audio file.
[0131] For example, taking the target audio file as a standard digital audio file in WAV format, WAV format is easier to preserve the quality of the original audio and is a lossless music format. After obtaining the initial format of the confidence list, the terminal device 200 can save the confidence list as a WAV format to ensure that the audio in the test audio file is not damaged. In actual operation, the confidence list can be saved as the target audio file using an audio conversion tool. For example, the confidence list can be input into an audio conversion tool, and after conversion, the audio conversion tool outputs the corresponding conversion result. The confidence list can also be saved as the target audio file in other ways; the specific method is not limited in this application. After step S4 is completed, step S5 can be executed.
[0132] Step S5: The terminal device 200 extracts the audio wave from the test audio file and imports the target audio file into the first audio track to obtain the confidence waveform of the target audio file.
[0133] To obtain the audio features in the test audio file, the terminal device 200 can extract audio waves from the test audio file. When extracting audio waves from the test audio file, the terminal device 200 first parses the test audio file to obtain the audio amplitude information of the test audio file, then iterates through the audio amplitude information to obtain the vibration range of the audio amplitude information, and then generates the audio waves of the test audio file based on the vibration range.
[0134] To obtain the audio features from the target audio file, the terminal device 200 can import the target audio file into the first audio track to obtain the confidence waveform of the target audio file. Since the target audio file is in a common audio format, its confidence waveform can be visually presented using audio analysis software. This allows users to intuitively see the confidence waveform effect of the target audio file, eliminating the need to analyze and summarize patterns from numerous confidence results individually. This solves the problem of unintuitive results when performing data analysis using model-based methods and also improves the processing efficiency of the terminal device 200. After step S5 is completed, step S6 can be executed.
[0135] Step S6: The terminal device 200 outputs the audio sound wave and confidence waveform in a visual form.
[0136] After extracting the audio waveform from the test audio file and importing the target audio file into the first audio track to obtain the confidence waveform of the target audio file, the terminal device 200 can output the audio waveform and confidence waveform in a visual form. The terminal device 200 first obtains the display channel where the confidence waveform is located, i.e., the display channel of the first audio track. Then, it creates a second audio track in this display channel and imports the audio waveform into the second audio track. In this way, the first audio track containing the confidence waveform of the target audio file and the second audio track containing the audio waveform of the test audio file will be displayed in the same display channel. This allows for the simultaneous display of the audio waveform and confidence waveform in different audio tracks within the same display channel, i.e., simultaneous visual output of the audio waveform and confidence waveform. Thus, by displaying the confidence waveform in the first audio track and the audio waveform in the second audio track within the same display channel, the characteristics of the audio data can be intuitively understood through the confidence waveform of the target audio file and the audio waveform of the test audio file. After step S6 is completed, step S7 can be executed.
[0137] Step S7: The terminal device 200 compares the audio sound wave with the confidence waveform, generates waveform difference data, and generates analysis results of the test audio file based on the waveform difference data.
[0138] In order to obtain the difference between the audio sound wave and the confidence waveform to identify problematic audio data, in some embodiments, after obtaining the confidence waveform of the target audio file and the audio sound wave in the test audio file from step S6, the terminal device 200 compares the audio sound wave and the confidence waveform to generate waveform difference data, so as to generate the analysis result of the test audio file based on the waveform difference data, that is, to determine the problematic audio data based on the waveform difference data.
[0139] When generating waveform difference data, the terminal device 200 first iterates through the audio amplitude information of the test audio file, then iterates through the confidence waveform to obtain its amplitude information, and finally compares the audio amplitude information with the waveform amplitude information to obtain waveform difference data. For example, when iterating through the audio amplitude information of the test audio file, the terminal device 200 can obtain the audio amplitude range of the test audio file; when iterating through the confidence waveform, it can obtain the vibration range of the confidence waveform. Then, it can compare the audio amplitude range of the test audio file with the vibration range of the confidence waveform. Any discrepancies in the ranges or vibration amplitudes between the two can be identified as waveform difference data.
[0140] To enable users to intuitively see the location of waveform difference data, in some embodiments, the terminal device 200 can mark the range of waveform difference data. The terminal device 200 can first locate the waveform difference data, that is, obtain the first display position where the audio amplitude information of the test audio file is located and the second display position where the waveform amplitude information of the confidence waveform is located. Since both are in the same display channel, the first and second display positions can be merged to generate the total display position of the waveform difference data. This total display position is then marked in the display channel, thus achieving the goal of marking the total display position of the waveform difference data. In this way, by comparing the audio sound wave and the confidence waveform to generate waveform difference data, it is possible to pinpoint the specific part of the audio data in the test audio file that has a problem, rather than simply reflecting the overall characteristics of the audio. This improves the accuracy of the test audio file analysis and solves the problem of inaccurate results when performing data analysis using model analysis methods.
[0141] As can be seen from the above technical solutions, the above embodiments provide an audio data analysis method. The method aligns the analysis results of the test audio file and the preset model in the time domain by sampling rate, saves the confidence list as the target audio file, and outputs the analysis results in a visual form. This can solve the problem that the results are not intuitive and inaccurate when performing data analysis through model analysis methods.
[0142] The same or similar parts among the various embodiments in this specification can be referred to mutually, and will not be repeated here.
[0143] Those skilled in the art will clearly understand that the techniques in the embodiments of the present invention can be implemented using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions in the embodiments of the present invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or certain parts of the embodiments of the present invention.
[0144] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
[0145] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, thereby enabling those skilled in the art to better utilize the described embodiments and various different variations of embodiments suitable for specific use considerations.
Claims
1. A terminal device, characterized in that, include: A memory, wherein program instructions are stored; The processor, by executing the program instructions, is configured to: Input the test audio file into the preset model; Obtain the confidence score result obtained by the preset model performing inference on the test audio file, and store the confidence score result as a confidence score list; The sampling rate of the test audio file is calculated according to the inference step size of the preset model; The confidence list is saved as a target audio file based on the sampling rate. The audio wave is extracted from the test audio file, and the target audio file is imported into the first audio track to obtain the confidence waveform of the target audio file; The audio wave and the confidence waveform are output in a visual form; By comparing the audio sound wave with the confidence waveform, waveform difference data is generated, and analysis results of the test audio file are generated based on the waveform difference data.
2. The terminal device according to claim 1, characterized in that, The step of the processor executing the step of obtaining the confidence result obtained by the preset model performing inference on the test audio file is further configured to: Generate a request instruction to obtain the confidence score result; In response to the request instruction, receive the confidence result returned by the preset model according to the request instruction; Obtain the output format of the confidence score result; Output the confidence score results according to the specified output format.
3. The terminal device according to claim 2, characterized in that, The processor performing the step of storing the confidence results as a confidence list is further configured to: Iterate through the confidence scores; A confidence list is created based on the output format to store the confidence results; Add the confidence score result to the confidence score list.
4. The terminal device according to claim 1, characterized in that, The processor, which performs the step of calculating the sampling rate of the test audio file according to the inference step size of the preset model, is further configured to: Obtain the unit time used to sample the test audio file; Obtain the inference step size of the preset model; The sampling rate of the test audio file is calculated based on the ratio of the unit time to the inference step size.
5. The terminal device according to claim 1, characterized in that, The processor performing the step of saving the confidence list as a target audio file based on the sampling rate is further configured to: Iterate through the confidence list to obtain its initial format; Obtain the target audio format of the target audio file; The initial format is converted into the target audio format so that the confidence list is saved as the target audio file.
6. The terminal device according to claim 1, characterized in that, The processor performs the step of extracting audio waves from the test audio file, and is further configured to: The test audio file is parsed to obtain the audio amplitude information of the test audio file; Traverse the audio amplitude information to obtain the vibration range of the audio amplitude information; The audio sound waves of the test audio file are generated based on the vibration range.
7. The terminal device according to claim 6, characterized in that, The processor performing the step of outputting the audio sound wave and the confidence waveform in a visual form is further configured to: Obtain the display channel where the confidence waveform is located; Create a second audio track in the display channel; The audio wave is imported into the second audio track to simultaneously display the audio wave and the confidence waveform in the same display channel but on different audio tracks.
8. The terminal device according to claim 7, characterized in that, The step of the processor performing the comparison of the audio sound wave and the confidence waveform to generate waveform difference data is further configured to: Iterate through the audio amplitude information of the test audio file; By traversing the confidence waveform, the waveform amplitude information of the confidence waveform is obtained; By comparing the audio amplitude information and the waveform amplitude information, waveform difference data is obtained.
9. The terminal device according to claim 8, characterized in that, The processor is further configured to: Locate the waveform difference data to obtain the first display position where the audio amplitude information is located and the second display position where the waveform amplitude information is located; The first display position and the second display position are merged to generate the total display position of the waveform difference data; Mark the total display position in the display channel.
10. A method for analyzing audio data, applied to a terminal device, characterized in that, include: Input the test audio file into the preset model; Obtain the confidence score result obtained by the preset model performing inference on the test audio file, and store the confidence score result as a confidence score list; The sampling rate of the test audio file is calculated according to the inference step size of the preset model; The confidence list is saved as a target audio file based on the sampling rate. The audio wave is extracted from the test audio file, and the target audio file is imported into the first audio track to obtain the confidence waveform of the target audio file; The audio wave and the confidence waveform are output in a visual form; By comparing the audio sound wave with the confidence waveform, waveform difference data is generated, and analysis results of the test audio file are generated based on the waveform difference data.
Citation Information
Patent Citations
Data processing method and equipment
CN113506584A
Classification model training method and device, computer equipment and storage medium
CN115240659A