Audio detection method and device, electronic equipment and computer readable storage medium

By obtaining and comparing the similarity of the spectrum graph and frequency domain graph of the source audio file and the recorded audio file, the problem of insufficient audio detection accuracy in gaming scenarios is solved, and the accuracy of audio detection and user experience are improved.

CN120853609APending Publication Date: 2025-10-28TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410531181.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-28
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing technologies lack accuracy in audio detection in gaming scenarios, making it difficult to effectively match multiple audio features, resulting in a reduced gaming experience.

Method used

By obtaining the spectrum graph and frequency domain graph of the source audio file and the recorded audio file, calculating the similarity between the two, determining the detection result of the audio file, and combining the similarity of the spectrum graph and frequency domain graph to perform audio detection.

Benefits of technology

It improves the accuracy of audio detection, ensures the correctness and timeliness of audio files, and enhances the user experience in game scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853609A_ABST
    Figure CN120853609A_ABST
Patent Text Reader

Abstract

The invention provides an audio detection method and device, electronic equipment and a computer readable storage medium. The method comprises the steps that a first audio file and a to-be-detected second audio file are acquired, the first audio file is a source audio file, and the second audio file is recorded based on the first audio file and is based on a first spectrogram of the first audio file and a second spectrogram of the second audio file; determining a first similarity between the first audio file and the second audio file; and determining a second similarity between the first audio file and the second audio file based on the first frequency domain graph of the first audio file and the second frequency domain graph of the second audio file, and determining a detection result corresponding to the second audio file based on the first similarity and the second similarity. According to the invention, the accuracy of audio detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to audio detection technology, and more particularly to an audio detection method, apparatus, electronic device, and computer-readable storage medium. Background Technology

[0002] Audio testing in game scenarios primarily relies on manual verification. Based on human gaming experience, it depends on auditory assessment of audio accuracy and latency to intuitively judge playback performance. In related technologies, most audio matching algorithms are typically based on human voice characteristics for voice matching. However, unlike the singular characteristics of the human voice, game audio features are much richer, making voice-based matching algorithms less effective in game scenarios.

[0003] Currently, there is no good method to improve the accuracy of audio detection in game scenarios. Summary of the Invention

[0004] This application provides an audio detection method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can improve the accuracy of audio detection.

[0005] The technical solution of this application embodiment is implemented as follows:

[0006] This application provides an audio detection method, the method comprising:

[0007] Obtain a first audio file and a second audio file to be detected, wherein the first audio file is the source audio file and the second audio file is recorded based on the first audio file;

[0008] Based on the first spectrogram of the first audio file and the second spectrogram of the second audio file, a first similarity between the first audio file and the second audio file is determined.

[0009] Based on the first frequency domain diagram of the first audio file and the second frequency domain diagram of the second audio file, a second similarity between the first audio file and the second audio file is determined.

[0010] Based on the first similarity and the second similarity, the detection result corresponding to the second audio file is determined.

[0011] This application provides an audio detection device, including:

[0012] An audio acquisition module is used to acquire a first audio file and a second audio file to be detected, wherein the first audio file is a source audio file and the second audio file is recorded based on the first audio file;

[0013] The similarity matching module is used to determine a first similarity between the first audio file and the second audio file based on a first spectrogram of the first audio file and a second spectrogram of the second audio file; and to determine a second similarity between the first audio file and the second audio file based on a first frequency domain diagram of the first audio file and a second frequency domain diagram of the second audio file.

[0014] The detection result determination module is used to determine the detection result corresponding to the second audio file based on the first similarity and the second similarity.

[0015] This application provides an electronic device, including:

[0016] Memory is used to store executable instructions for a computer;

[0017] The processor, when executing computer-executable instructions stored in the memory, implements the audio detection method provided in the embodiments of this application.

[0018] This application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the audio detection method provided in this application when executed by a processor.

[0019] This application provides a computer program product, including a computer program or computer executable instructions, which, when executed by a processor, implements the audio detection method provided in this application.

[0020] The embodiments of this application have the following beneficial effects:

[0021] The system records a source audio file and obtains its spectrogram and frequency domain graph for both. By comparing the similarity between the spectrograms and frequency domain graphs of the source and recorded audio files, the system can determine the correctness of the played audio based on multiple audio features. By combining the similarity between the spectrograms and the frequency domain graphs, the system can determine the detection result corresponding to the recorded audio file, thereby improving the accuracy of audio detection. Attached Figure Description

[0022] Figure 1A This is a first schematic diagram illustrating the application mode of the audio detection method provided in the embodiments of this application;

[0023] Figure 1B This is a second schematic diagram illustrating the application mode of the audio detection method provided in the embodiments of this application;

[0024] Figure 2A This is a schematic diagram of the structure of the terminal device provided in the embodiments of this application;

[0025] Figure 2B This is a schematic diagram of the server structure provided in an embodiment of this application;

[0026] Figure 3A This is a first flowchart illustrating the audio detection method provided in this application embodiment;

[0027] Figure 3B This is a schematic diagram of the second process of the audio detection method provided in the embodiments of this application;

[0028] Figure 3C This is a schematic diagram of the third process of the audio detection method provided in the embodiments of this application;

[0029] Figure 3D This is a schematic diagram of the fourth process of the audio detection method provided in the embodiments of this application;

[0030] Figure 3E This is a schematic diagram of the fifth process of the audio detection method provided in the embodiments of this application;

[0031] Figure 4 This is a schematic diagram of the sixth process of the audio detection method provided in the embodiments of this application;

[0032] Figure 5 This is a schematic diagram illustrating the principle of the audio detection method provided in the embodiments of this application;

[0033] Figure 6 This is a model structure diagram of the convolutional neural network provided in the embodiments of this application;

[0034] Figure 7 This is a schematic diagram illustrating the inclusion relationship between the sub-audio and audio source file frequency domain diagrams provided in the embodiments of this application. Detailed Implementation

[0035] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0036] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0037] If the application documents contain similar descriptions such as "first / second", the following explanation shall be added: In the following description, the terms "first / second / third" are used only to distinguish similar objects and do not represent a specific order of objects. It is understood that "first / second / third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0038] It is understood that, in the embodiments of this application, the collection and processing of relevant data (e.g., background music, system prompts, and voice prompts in game programs) should be strictly carried out in accordance with the requirements of relevant national laws and regulations, obtaining the informed consent or separate consent of the personal information subject, and conducting subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0039] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0040] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.

[0041] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant national laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0042] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0043] 1) Spectrum diagram: A graphical representation of a signal at various frequencies, using a wave pattern on the horizontal and vertical axes. It is a type of graph that describes the spectral information of a signal on a time-frequency plane. Spectrum diagrams typically use time as the horizontal axis and frequency as the vertical axis, using different colors or grayscale values ​​to represent the intensity of different frequency components.

[0044] 2) Frequency domain graph: This graph shows the frequency distribution of a signal in the frequency domain. It is usually the result of a Fourier transform of the signal and is plotted on a one-dimensional graph. The horizontal axis represents frequency and the vertical axis represents the amplitude or power of the signal. It is mainly used to display the frequency components of the signal and can clearly show the frequency components and relative intensity of the signal. It only displays the frequency domain information of the signal and does not include time domain (time) information.

[0045] 3) Fast Forward Moving Picture Experts Group (FFMPEG): This is an open-source audio and video processing software that can perform recording, conversion, encoding, and decoding operations.

[0046] 4) Proxy method: It is a protective layer for creating objects, also known as an object wrapper. It allows the creation of a special object that can intercept all method calls of other objects. Based on the created object, it is possible to record the time when the audio function is called, the parameters called, etc., and to view this information at a later point in time.

[0047] 5) Absolute Time: This is a fixed, universally accepted time reference point used for time stamping and sorting audio events. It is based on an external clock or calendar system, such as Greenwich Mean Time (GMT) or Coordinated Universal Time (UTC). Using absolute time ensures that the recording and analysis of audio events are consistent and unaffected by other factors (such as volume, sound effects, etc.).

[0048] 6) Fast Fourier Transform (FFT): Used to calculate the Fourier transform of discrete signals. The Fast Fourier Transform (FFT) can convert time-domain signals into frequency-domain signals, and is suitable for processing periodic signals, noise signals, audio signals, etc. In digital signal processing, the Fast Fourier Transform is commonly used for tasks such as spectrum analysis, filtering, and feature extraction.

[0049] 7) Mini-games: These are game applications that can be used without downloading or installing. They reside within the software or platform, spread through the social attributes of the host software or platform, have a small file size, and offer a lightweight experience. Examples include mini-games that run as mini-programs embedded in any app, or mini-games that can be played simply by downloading them to a browser.

[0050] Audio matching in related technologies typically involves detecting human voices. However, the types of audio in game scenarios are much richer and differ from the single characteristic of human voices. Matching audio in game scenarios requires consideration of multiple audio features, and matching algorithms in related technologies often fail to achieve good results. This application provides an audio detection method, an audio detection device, an electronic device, a computer-readable storage medium, and a computer program product that can improve the accuracy of audio detection.

[0051] The following describes exemplary applications of the electronic devices provided in the embodiments of this application. These devices can be implemented as various types of terminals such as laptops, tablets, desktop computers, set-top boxes, smartphones, smart speakers, smartwatches, smart TVs, and in-vehicle terminals, or as servers. Exemplary applications when the device is implemented as a terminal or server will be described below.

[0052] In the Figure 1A Before proceeding, let's first introduce the game modes involved in the terminal device and server collaborative implementation scheme. This scheme primarily involves two game modes: local game mode and cloud game mode. In local game mode, the terminal device and server collaboratively run the game processing logic. The player's input commands on the terminal device are partly processed by the terminal device's game logic, and partly by the server. Furthermore, the server-side game logic processing is often more complex and requires more computing power. In cloud game mode, the server handles all game logic processing, and the cloud server renders the game scene data into audio and video streams, which are then transmitted to the terminal device for display over the network. The terminal device only needs basic streaming media playback capabilities and the ability to receive player commands and send them to the server.

[0053] See Figure 1A , Figure 1A This is a first schematic diagram illustrating the application mode of the audio detection method provided in this application embodiment. To support an audio detection application, an example is provided. Figure 1A The system involves server 200, network 300, terminal device 400 and database 500. Terminal device 400 is connected to server 200 through network 300. Network 300 can be a wide area network or a local area network, or a combination of both.

[0054] In some embodiments, the user can be an audio tester or a gamer, server 200 is a server used to test audio files, terminal device 400 is a user-operated terminal, and terminal device 400 has an application installed that can run games and record audio. Database 500 stores source audio and recorded audio files, and game application interface 100 is used to display the running game interface.

[0055] For example, terminal device 400 is used to receive user operation instructions and record audio in the game scene running on game application interface 100 according to the operation instructions. Terminal device 400 sends the recorded audio file and the source file (equivalent to the first audio file and the second audio file to be detected in this embodiment) together to server 200 through network 300. Server 200 matches the similarity of the spectrogram and frequency domain diagram of the source file and the recorded audio file to obtain the audio correctness detection result, and sends the audio correctness detection result to terminal device 400 through network 300. Terminal device 400 provides feedback on the audio correctness detection result to the user. At this time, testers or game players can correctly judge the audio playback situation in the current game scene through the feedback test result.

[0056] In some embodiments, the audio detection method of this application can also be applied in the following application scenarios:

[0057] 1. In map navigation application scenarios, the audio detection method provided in the embodiments of this application is called to detect the correctness of audio information in different scenarios in map navigation. For example, the background sound, driving prompt sound, road condition prompt sound, and remaining distance prompt sound in different map navigation scenarios are detected to ensure timely and accurate playback, thereby improving the user experience when using the navigation program and avoiding situations where route navigation is not timely or the route is incorrect due to the failure to play audio in a timely manner.

[0058] 2. In music game application scenarios, the audio detection method provided in the embodiments of this application is called to detect the accuracy and timeliness of background music and sound effects. For example, it can detect whether the drum beats in the music game are played correctly and in a timely manner, improve the matching degree between sound effects and drum beats and user operation actions, avoid the user's game experience being worsened due to wrong beats or delayed playback, miss scoring opportunities, improve the user's game experience, and enhance the fun and interactivity of the game.

[0059] 3. In the application scenario of instant messaging software, the audio detection method provided in the embodiments of this application can be used to detect whether various prompts in the instant messaging software are played in a timely manner, so as to ensure that users do not miss important messages during use. For example, the detection of whether the prompts for new messages, call prompts, file transfer and receiving completion prompts are played in a timely manner, and whether they are correct, can prevent users from failing to view important information or answer calls in a timely manner due to untimely or incorrect playback of prompts.

[0060] 4. In the application scenario of mini-games, the audio detection method provided in the embodiments of this application can be called to detect whether the background music, game sound effects, plot dialogue and real-time voice of the mini-game running on the host software or platform are played in a timely manner and whether they are played correctly. This prevents the interactive feedback in the mini-game from being out of sync, the gameplay from being missing, and the game content from being discussed in a timely manner due to the audio not being played or being delayed, which reduces the user experience and immersion of the mini-game.

[0061] This application embodiment can be implemented using database technology. A database, simply put, can be viewed as an electronic filing cabinet storing electronic files, where users can perform operations such as adding, querying, updating, and deleting data. A "database" is a collection of data stored together in a certain way, capable of being shared by multiple users, having minimal redundancy, and being independent of application programs.

[0062] A Database Management System (DBMS) is a computer software system designed to manage databases, generally possessing basic functions such as storage, retrieval, security, and backup. DBMSs can be classified according to the database model they support, such as relational or XML (Extensible Markup Language); or according to the type of computer they support, such as server clusters or mobile devices; or according to the query language used, such as Structured Query Language (SQL) or XQuery; or according to performance priorities, such as maximum scale or maximum operating speed; or other classification methods. Regardless of the classification method used, some DBMSs can cross categories, for example, simultaneously supporting multiple query languages.

[0063] This application embodiment can also be implemented using cloud technology. Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, and application technology based on cloud computing business models. It can form a resource pool, available on demand, offering flexibility and convenience. Cloud computing technology will become a crucial support. Backend services of technical network systems require substantial computing and storage resources, such as video websites, image websites, and many portal websites. With the rapid development and application of the internet industry, and driven by demands for search services, social networks, mobile commerce, and open collaboration, every item may eventually possess its own hash-coded identification mark, requiring transmission to a backend system for logical processing. Data at different levels will be processed separately, and various industry data will require robust system support, which can only be achieved through cloud computing.

[0064] This application embodiment can also be understood through the theory, methods, technologies, and application systems of Artificial Intelligence (AI), which utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, artificial intelligence is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine capable of reacting in a manner similar to human intelligence. Artificial intelligence, in essence, studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0065] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interactive systems, and mechatronics. Among these, pre-trained models, also known as large-scale models or foundational models, can be widely applied to downstream tasks across various AI fields after fine-tuning. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0066] In some embodiments, server 200 may be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Terminals and servers can be connected directly or indirectly via wired or wireless communication, which is not limited in this embodiment.

[0067] See Figure 1B , Figure 1B This is a second schematic diagram illustrating the application mode of the audio detection method provided in the embodiments of this application; for example, Figure 1B The invention involves a terminal device 400, which has an application installed that can run games and record audio. The terminal device 400 can display a game application interface 100, which is used to display the running game interface. The terminal device 400 stores source audio and recorded audio files.

[0068] In some embodiments, the user can be an audio tester or a game player. The terminal device 400 is a user-operated terminal used to receive user operation commands and record audio from the game scene running on the game application interface 100 according to the operation commands. The terminal device 400 performs audio detection based on the recorded audio file and the source file, obtains the audio correctness detection result, and feeds back the audio correctness detection result to the user. The user judges the audio playback status in the current game scene based on the feedback test result.

[0069] See Figure 2A , Figure 2A This is a schematic diagram of the structure of the terminal device provided in the embodiments of this application. Figure 2A The terminal device 400 shown includes at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. The various components in the terminal device 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to implement communication between these components. In addition to a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 2A The general labeled all buses as Bus System 440.

[0070] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0071] User interface 430 includes one or more output devices 431 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0072] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices physically located away from the processor 410.

[0073] The memory 450 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 450 described in this application embodiment is intended to include any suitable type of memory.

[0074] In some embodiments, memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0075] Operating system 451 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks;

[0076] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc.

[0077] Presentation module 453 is configured to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 431 (e.g., a display screen, a speaker, etc.) associated with user interface 430;

[0078] The input processing module 454 is used to detect and translate one or more user inputs or interactions from one or more input devices 432.

[0079] In some embodiments, the apparatus provided in this application can be implemented in software. Figure 2A An audio detection device 455 stored in memory 450 is shown. This device can be software in the form of programs and plugins, including the following software modules: an audio recording module 4551, a similarity matching module 4552, and a detection result determination module 4553. These modules are logically linked and can therefore be arbitrarily combined or further separated according to their implemented functions. Figure 2A For ease of explanation, all the above modules are shown at once, but this should not be interpreted as excluding the implementation of the audio detection device 455 which may only include the audio recording module 4551. The functions of each module will be explained below.

[0080] See Figure 2B , Figure 2B This is a schematic diagram of the server structure provided in an embodiment of this application. Figure 2B The server 200 shown includes at least one processor 210, memory 250, and at least one network interface 220. The various components of server 200 are coupled together via a bus system 240. It is understood that the bus system 240 is used to implement communication between these components. In addition to a data bus, the bus system 240 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 2B The general labeled all buses as Bus System 240.

[0081] Processor 210 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0082] The memory 250 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 250 may optionally include one or more storage devices physically located away from the processor 210.

[0083] Memory 250 may include volatile memory or non-volatile memory, or both. Non-volatile memory may be read-only memory (ROM), and volatile memory may be random access memory (RAM). The memory 250 described in this application embodiment is intended to include any suitable type of memory.

[0084] In some embodiments, memory 250 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0085] Operating system 251 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks;

[0086] The network communication module 252 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 220, exemplary network interfaces 220 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc.

[0087] In some embodiments, the apparatus provided in this application can be implemented in software. Figure 2B An audio detection device 255 stored in memory 250 is shown. This device can be software in the form of programs and plugins, including the following software modules: an audio recording module 2551, a similarity matching module 2552, and a detection result determination module 2553. These modules are logically linked and can therefore be arbitrarily combined or further separated according to their implemented functions. Figure 2B For ease of explanation, all the above modules are shown at once, but this should not be interpreted as excluding the implementation of the audio detection device 255 which may only include the audio recording module 2551. The functions of each module will be explained below.

[0088] In some embodiments, the terminal or server can implement the audio detection method provided in this application by running various computer-executable instructions or computer programs. For example, computer-executable instructions can be microprogram-level commands, machine instructions, or software instructions. Computer programs can be native programs or software modules in an operating system; they can be native applications (APPs), i.e., programs that need to be installed in the operating system to run, such as game APPs or map navigation APPs; or they can be applets that can be embedded in any APP, i.e., programs that only need to be downloaded to a browser environment to run. In summary, the aforementioned computer-executable instructions can be any form of instruction, and the aforementioned computer programs can be any form of application, module, or plugin.

[0089] The audio detection method provided in this application will be described in conjunction with exemplary applications and implementations of the electronic devices provided in the embodiments of this application.

[0090] The audio detection method provided in the embodiments of this application will be described below. As mentioned above, the electronic device implementing the image processing method of the embodiments of this application can be a terminal, a server, or a combination of both. Therefore, the executing entity of each step will not be described again below.

[0091] See Figure 3A , Figure 3A This is a first flowchart illustrating the audio detection method provided in this application embodiment, which will be combined with... Figure 3A The steps shown are explained.

[0092] In step 301, the first audio file and the second audio file to be detected are obtained.

[0093] Here, the first audio file is the source audio file, and the second audio file is recorded based on the first audio file.

[0094] For example, to detect whether audio is playing normally in different application scenarios, it is necessary to record the entire audio playback process. Recording is performed while the source audio file is playing to obtain the audio file to be detected. In this embodiment of the application, the background music playback of a game application is used as an example for explanation. The audio detection method provided in this embodiment of the application can also be applied to application scenarios such as detecting the prompts of map applications and the prompts of instant messaging software.

[0095] See Figure 5 , Figure 5 This is a schematic diagram illustrating the principle of the audio detection method provided in the embodiments of this application; Figure 5 When the game interface 510 is running, it plays different types of audio. At this time, the recording control in the terminal device is triggered to record the first audio file being played. The recording method is screen recording of the terminal device, that is, without the need for the terminal device's microphone or other recording devices, it directly calls the application's internal interface to record the audio being played, and obtains the second audio file to be detected.

[0096] In some embodiments, see Figure 3B , Figure 3B This is a schematic diagram of the second process of the audio detection method provided in the embodiments of this application; to illustrate the audio recording process in more detail, during the execution... Figure 3A Before step 301, execute Figure 3B Steps 3011 to 3012 are used to obtain the second audio file, as detailed below.

[0097] In step 3011, a virtual scene is displayed in the human-computer interaction interface.

[0098] Here, the first audio file is played when the virtual scene runs.

[0099] For example, a virtual scene can be a virtual scene in a game, a virtual scene in a social application, or a virtual scene in a map application that is mapped from the real world.

[0100] For example, the playback of the first audio file relies on a human-computer interaction interface (HCI). An HCI is the interface through which people communicate and interact with computers or other terminal devices, typically including a graphical user interface (GUI) and voice recognition. Users communicate and operate the terminal device through commands, triggering different virtual scenarios. The HCI displays the virtual scenarios of the software program under test, and the display and operation of these virtual scenarios are accompanied by different types of audio playback, providing an immersive user experience.

[0101] In some embodiments, the human-computer interaction interface runs through a mini-program embedded in the application, the virtual scene is a mini-game virtual scene, and recording is achieved through the application's audio output interface; displaying the virtual scene in the human-computer interaction interface includes: in response to a trigger operation on the mini-program in the application, displaying the mini-game virtual scene associated with the mini-program.

[0102] For example, the audio recording of the mini-game is implemented through the application's internal interface. Compared with the recording method of recording through the microphone of the terminal device in related technologies, the recording efficiency is higher and the response speed of the recording process is faster, which improves the accuracy of the recorded file.

[0103] For example, mini-games reside within the software or platform and are distributed through the host software or platform. Therefore, mini-games are small in size and require no additional download or installation. They do not have a separate icon on electronic devices, are not limited by the operating system or terminal device, and offer a high degree of user interaction. (Continue to refer to...) Figure 5 , Figure 5 The game interface 510 in the middle displays the virtual game scene displayed in the human-computer interaction interface. The game interface 510 can be the interface for a mini game, which can be run by a mini program embedded in instant messaging software.

[0104] In step 3012, in response to a trigger operation on the recording control, the first audio file being played is recorded to obtain a second audio file.

[0105] For example, a recording control is a control in a screen recording application on a terminal device. It is used to trigger the recording of the screen displayed and the audio played in the human-computer interaction interface. The triggering operation can be to start or stop recording by controlling the recording control. The type of triggering operation can be click, long press, gesture, etc. The recording control is provided by the computer or other terminal device running the application itself. During the recording process, the recording time is displayed, and after the recording is completed, the content of the recorded audio file can be viewed.

[0106] While the application is running, the recording control is triggered to record the audio information being played. The audio information can be recorded and saved, and the saved audio file can be played, edited, or shared.

[0107] In some embodiments, step 3012 can be implemented by the following method: when the virtual scene is running and the first audio file is in playback state, the first audio file being played is recorded to obtain a recording file, and the audio output interface is called to determine the playback start time, playback duration and playback end time of the first audio file; based on the playback start time, the playback duration and the playback end time, the recording file is trimmed to obtain at least one second audio file.

[0108] For example, when a virtual scene is running and an audio file is playing, the audio file is recorded. During recording, the audio output interface is called to determine the start time, end time, and duration of the audio playback. The audio output interface is provided by the computer or other terminal device platform running the virtual scene. It manages and controls the audio data through a proxy class, using proxy methods provided by JavaScript. These proxy methods can record the time and parameters of method calls to obtain audio call information. During recording, audio recording and obtaining audio call information through the interface are executed synchronously. This relies on the interface provided by the software or hardware environment that provides audio playback functionality. This interface is provided by an audio library written in JavaScript and can process the audio data.

[0109] The information obtained by calling the audio output interface may include: the storage path of the audio file, the start time and end time of the audio file playback, the playback duration, the duration of the audio playback, and the audio playback rate.

[0110] For example, after receiving the audio information from the API, the recorded audio file is trimmed based on the start time, duration, and end time of playback, resulting in at least one trimmed second audio file. The first audio file can be played in a loop, with one second audio file generated during each loop, the number of second audio files equal to the number of loops. Alternatively, the first audio file can be played only once, with the second audio file being a recording of the entire playback process of the first audio file.

[0111] In some embodiments, before performing step 3012, the following system times are also calibrated: the first system time of the human-computer interaction interface, the second system time of the virtual scene, and the third system time of the audio output interface corresponding to the first audio file.

[0112] For example, the first system time is the time when the user operates the computer or other terminal device, the second system time is the system time in the running application software, and the third system time is the time of the recording interface system.

[0113] At the start of recording, all system times are calibrated to ensure consistency. Audio file trimming is also based on absolute time, which is a fixed and universally accepted time reference point used for time stamping and sorting audio events. Based on an external clock or calendar system, using absolute time ensures that the recording and analysis of audio files are consistent and unaffected by volume or sound effects.

[0114] In this embodiment, the audio being played is recorded, and all system times are calibrated at the start of recording. The recorded audio is then trimmed based on absolute time, which maintains the consistency of audio file recording and analysis in terms of time, and is not affected by other factors, making the information obtained by calling the audio output interface more accurate.

[0115] Continue to refer to Figure 3A In step 302, a first similarity between the first audio file and the second audio file is determined based on the first spectrogram of the first audio file and the second spectrogram of the second audio file.

[0116] For example, after recording an audio file, a similarity comparison is performed between the first and second audio files. This is done by comparing the similarity between the source and recorded audio files' corresponding spectrograms. Based on the similarity comparison results, the inclusion relationship between the first and second audio files is determined. A spectrogram records the waveform of a signal at different frequencies; it's an image describing the signal's spectral information on a plane with time as the horizontal axis and frequency as the vertical axis. The intensity of different frequency components can be seen from the spectrogram. The inclusion relationship is the degree of similarity between two spectrograms; the numerical value of the similarity between the spectrograms serves as the inclusion coefficient, characterizing the inclusion relationship between the two images.

[0117] For ease of understanding, the following is combined with Figure 5 Explain the inclusion relationship, for Figure 5 The similarity of the cropped sub-audio file spectrum image 522 and the partially recorded source audio file spectrum image 521 is matched. It can be seen that there are partially similar image features between the cropped sub-audio file spectrum image 522 and the partially recorded source audio file spectrum image 521. Therefore, there is an inclusion relationship between the two spectrum images. The similarity of the spectrum images is a value between 0 and 1, which serves as the inclusion coefficient representing the inclusion relationship. For example, when the inclusion coefficient is 1, the inclusion relationship between the two spectrum images is complete, indicating that one spectrum image is completely contained within the other; when the inclusion coefficient is 0, the inclusion relationship between the two spectrum images is no inclusion, indicating that one spectrum image and the other spectrum image have no overlapping parts.

[0118] Continue to refer to Figure 5 ,from Figure 5 A portion of the recorded source audio file spectrum image 521 can be extracted from the source audio file spectrum image 520 and matched with the cropped sub-audio file spectrum image 522 to obtain the spectrum similarity matching result of the audio file.

[0119] In some embodiments, see Figure 3C , Figure 3C This is a schematic diagram of the third process of the audio detection method provided in the embodiments of this application; step 302 can be implemented by steps 3021 to 3023, as described in detail below.

[0120] In step 3021, a first spectrogram of the first audio file is determined based on the frequencies of the first audio file at different times.

[0121] For example, based on the frequency information of the first audio file at different times, a corresponding spectrogram is plotted for the first audio file. The spectrogram of the first audio file shows the frequency change of the source audio over time. The horizontal axis of the spectrogram represents time, and the vertical axis represents frequency.

[0122] In step 3022, a second spectrogram of the second audio file is determined based on the frequencies of the second audio file at different times.

[0123] For example, based on the frequency information of the recorded second audio file at different times, a corresponding spectrogram is plotted for the second audio file, showing the frequency changes of the recorded audio over time.

[0124] In step 3023, the first spectrogram and the second spectrogram are matched to obtain the first similarity between the first audio file and the second audio file.

[0125] For example, after obtaining the spectrograms corresponding to the first and second audio files, a similarity matching process is performed on the two spectrograms to obtain their image similarity. Similarity matching can be achieved by training a convolutional neural network model using an image similarity matching algorithm.

[0126] In some embodiments, step 3023 can be implemented by the following method: normalizing the first spectrogram and the second spectrogram respectively to obtain a normalized first spectrogram and a normalized second spectrogram; performing feature extraction processing on the normalized first spectrogram to obtain a first image feature; performing feature extraction processing on the normalized second spectrogram to obtain a second image feature; determining a third similarity between the first image feature and the second image feature; and using the third similarity as a first similarity between the first audio file and the second audio file.

[0127] In some embodiments, in the similarity matching process of spectrograms, normalization processing, feature extraction processing, and the processing of determining the third similarity are implemented by a convolutional neural network model. When the first image feature and the second image feature are represented as feature vectors, the third similarity is the cosine similarity between the first image feature and the second image feature.

[0128] The following sections will explain the convolutional neural network model and the method for determining the third similarity.

[0129] See Figure 6 , Figure 6 This is a structural diagram of the convolutional neural network model provided in the embodiments of this application; through Figure 6 The convolutional neural network model 601 in the text can perform the similarity matching processing of the spectrogram in step 3023, as explained in detail below.

[0130] For example, the normalization layer 6011 is used to normalize the spectrogram image in the input convolutional neural network model 601 to obtain the normalized first spectrogram and the second spectrogram. The normalization process includes: scaling or cropping the first spectrogram and the second spectrogram, adjusting the image size to a uniform resolution to ensure that the images have the same size, calculating the mean of the colors, and normalizing the pixel values ​​of the image to the [0, 1] interval. The normalized spectrograms have the same size and color space. The feature extraction layer 6012 performs convolution and pooling operations on the input image to extract image features from the two spectrograms, resulting in image features extracted from the normalized spectrograms. After extracting the image features of the two spectrograms, the feature comparison layer 6013 compares the extracted image features. Feature comparison can be achieved using the cosine similarity or Euclidean distance method of the feature vectors. The evaluation layer 6014 calculates the similarity between the two spectrograms based on the feature comparison results. The similarity can be a value between 0 and 1, representing the degree of similarity between the two images. That is, it outputs the similarity matching result of the two spectrograms, using the similarity value as the inclusion coefficient characterizing the inclusion relationship between the two images.

[0131] For example, when the first image feature and the second image feature are represented as feature vectors, cosine similarity can be used as a third similarity. Cosine similarity measures the cosine of the angle between two vectors, without considering the absolute magnitude of the two vectors. When the angle between two vectors is close to 0 degrees, the cosine similarity is close to 1, indicating that they are very similar in direction; when the angle is 90 degrees, the cosine similarity is 0, indicating that they are completely dissimilar in direction; when the angle is close to 180 degrees, the cosine similarity is close to -1, indicating that they are completely opposite in direction. Cosine similarity can effectively compare the similarity between feature vectors.

[0132] In this embodiment of the application, when performing similarity matching on the spectrogram, a convolutional neural network is trained and an image similarity matching algorithm is called for processing, which can more accurately obtain the degree of similarity matching. Spectrogram-based matching can more clearly reflect the frequency changes over time, providing a basis for adjusting audio playback parameters based on the similarity matching results.

[0133] Continue to refer to Figure 3A In step 303, a second similarity between the first audio file and the second audio file is determined based on the first frequency domain map of the first audio file and the second frequency domain map of the second audio file.

[0134] For example, after recording the audio files, a similarity comparison is performed between the first and second audio files based on their spectrograms. By comparing their respective frequency domain graphs, the frequency domain similarity between the first and second audio files is obtained. A frequency domain graph represents the frequency distribution of a signal in the frequency domain; the horizontal axis represents frequency, and the vertical axis represents the signal amplitude or power. By performing a Fourier transform on the signal, the frequency domain information is displayed more clearly. Based on the comparison results of the frequency domain graph similarity, the inclusion relationship between the source audio file and the recorded audio file can be determined more accurately. See also... Figure 7 , Figure 7 This is a schematic diagram illustrating the inclusion relationship between the sub-audio and audio source file frequency domain diagrams provided in the embodiments of this application; Figure 7 Region 7011 in the text represents the recorded sub-audio 701, and the region 7021 of the source audio 702 is included in the recording. Figure 7 It can be seen that the frequency domain distribution of region 7021 is contained within the frequency domain distribution of region 7011.

[0135] In some embodiments, see Figure 3D , Figure 3D This is a schematic diagram of the fourth process of the audio detection method provided in this application embodiment; step 303 can be achieved through... Figure 3D Steps 3031 to 3034 are implemented, and the details are explained below.

[0136] In step 3031, the audio time-domain signal of the first audio file is subjected to Fourier transform to obtain the first frequency-domain signal.

[0137] For example, based on the frequency distribution information of the first audio file, the signal of the first audio file is transformed using the Fast Fourier Transform (FFT) method to convert the time-domain signal into the corresponding frequency-domain signal. The FFT method decomposes the signal sequence into frequency and phase information, enabling frequency domain analysis of the signal. The horizontal axis of the frequency distribution curve represents frequency, and the vertical axis represents the amplitude or power of the signal.

[0138] In step 3032, the audio time-domain signal of the second audio file is subjected to Fourier transform to obtain the second frequency-domain signal.

[0139] For example, based on the frequency distribution information of the recorded second audio file, the signal of the second audio file is transformed by using the Fast Fourier Transform method to convert the time domain signal to obtain the frequency domain signal corresponding to the recorded audio.

[0140] In step 3033, a first frequency domain diagram is drawn based on the first frequency domain signal, and a second frequency domain diagram is drawn based on the second frequency domain signal.

[0141] Here, the horizontal axis of the first frequency domain diagram and the second frequency domain diagram represents frequency, and the vertical axis represents signal strength.

[0142] For example, based on the frequency information distribution of a first audio file, a corresponding frequency domain diagram is plotted for the first audio file. The distribution curve in the frequency domain diagram of the first audio file shows the frequency distribution of the source audio signal in the frequency domain. Based on the frequency information distribution of a second audio file, a corresponding frequency domain diagram is plotted for the second audio file. The distribution curve in the frequency domain diagram of the second audio file shows the frequency distribution of the recorded audio signal in the frequency domain.

[0143] In step 3034, a second similarity between the first audio file and the second audio file is determined based on the difference in signal intensity distribution between the first frequency domain map and the second frequency domain map.

[0144] For example, after obtaining two frequency domain maps, the distribution of signal intensity in the frequency domain maps is calculated. By calculating the difference in signal intensity distribution, the similarity between the first audio file and the second audio file can be determined.

[0145] In some embodiments, step 3034 can be implemented by performing the following process for each frequency that is the same in the first frequency domain map and the second frequency domain map: determining the signal strength difference between the second signal strength of the frequency in the second frequency domain map and the first signal strength in the first frequency domain map; counting a first number of positive signal strength differences; and taking the ratio between the first number and the total number of signal strength differences as the second similarity between the first audio file and the second audio file.

[0146] For example, for the same frequency, if the difference between the corresponding signal strengths in the two frequency domain graphs is negative, it means that the source audio exceeds the range of the recorded audio; if the difference between the corresponding signal strengths in the two frequency domain graphs is positive, it means that the recorded audio can cover the source audio. Finally, the percentage of coverage of the source audio by the recorded audio is calculated, and the percentage of coverage is used as the inclusion coefficient of the frequency domain graph to determine the inclusion relationship between the source audio and the recorded audio.

[0147] For example, the inclusion relationship between frequency domain graphs is determined by the ratio between the number of sampling points under the curve of the second frequency domain graph of the second audio file and the total number of sampling points of the curve of the first frequency domain graph of the first audio file, that is, the ratio between the first number of positive signal strength differences and the total number of signal strength differences. The sampling point is the point on the frequency curve corresponding to each frequency in the frequency domain graph, and each sampling point corresponds to a different frequency.

[0148] Starting from frequency 0 to a maximum of 40,000 Hz (beyond 40,000 Hz, the human ear cannot distinguish the difference, so calculations are not performed), the input signal sequence is transformed using the Fast Fourier Transform (FFT) method. The difference between the recorded audio and the source audio distribution is calculated for each sampling point. For any given sampling point, a negative difference indicates that the source audio exceeds the recorded audio; a positive difference indicates that the recorded audio covers the source audio. For example, if the coverage coefficient of the frequency domain plot is greater than 80%, an inclusion relationship is determined; if the coverage coefficient is less than 80%, no inclusion relationship is determined.

[0149] See Figure 7 , Figure 7 This is a schematic diagram illustrating the inclusion relationship between the sub-audio and audio source file frequency domain diagrams provided in the embodiments of this application; Figure 7The sub-audio 701 recorded in the image represents the first frequency domain diagram of the present application embodiment, and has an inclusion relationship with the second frequency domain diagram of the present application embodiment represented by the source audio 702. Region 7011 represents the frequency and signal intensity distribution of the recorded sub-audio 701, and region 7021 represents the signal intensity distribution corresponding to the frequency of the source audio 702. The intensity distributions in region 7021 and region 7011 overlap. By calculating the inclusion coefficient in the two regions, it is possible to determine whether the inclusion relationship between the first frequency domain diagram and the second frequency domain diagram satisfies the frequency domain diagram inclusion threshold of the present application.

[0150] In this embodiment of the application, by comparing the frequency domain similarity between audio files, the signal strength partition corresponding to the frequency can be judged more accurately. The accurate value of similarity is obtained through a linear method, which provides a basis for adjusting the audio playback parameters according to the similarity matching results.

[0151] Continue to refer to Figure 3A In step 304, the detection result corresponding to the second audio file is determined based on the first similarity and the second similarity.

[0152] For example, after comparing the similarity of the spectrogram and frequency domain graph of the first audio file and the second audio file, the similarity comparison results of the spectrogram and the frequency domain graph can be obtained respectively. The similarity results of the spectrogram and the frequency domain graph are judged simultaneously. By judging whether the inclusion relationship of the two aspects is satisfied at the same time, the detection result corresponding to the audio file is determined.

[0153] In some embodiments, see Figure 3E , Figure 3E This is a schematic diagram of the fifth step of the audio detection method provided in this application embodiment; step 304 can be achieved through... Figure 3E Steps 3041 to 3042 are implemented, and the details are explained below.

[0154] In step 3041, in response to the first similarity and the second similarity satisfying the detection condition, the detection result corresponding to the second audio file is determined to be that the audio playback is normal.

[0155] Here, the detection conditions include: the first similarity is greater than the first similarity threshold, and the second similarity is greater than the second similarity threshold.

[0156] For example, when determining whether the similarity meets the detection conditions, the similarity of the spectrogram and the similarity of the frequency domain graph are judged simultaneously. When the similarity of both the spectrogram and the frequency domain graph is greater than the similarity threshold, it can be determined that the detection result corresponding to the audio file is that the audio playback is normal.

[0157] By analyzing the spectrogram similarity between the source audio file and the recorded audio file, the inclusion relationship between the two spectrograms can be determined. The inclusion relationships corresponding to the similarity values ​​are: complete inclusion (inclusion coefficient of 1), indicating that one spectrogram is completely contained within the other, meaning one spectrogram covers all parts of the other; partial overlap (inclusion coefficient close to 1), indicating that the two spectrograms partially overlap, meaning they have the same intensity at certain frequencies; and no inclusion (inclusion coefficient of 0), indicating that one spectrogram has no overlap with the other, meaning their frequency ranges and intensity distributions do not intersect.

[0158] After the similarity of the spectrograms is processed by the convolutional neural network model, a numerical result between 0 and 1 is obtained. This application sets the similarity threshold of the spectrograms to 0.8, that is, when the numerical value of the spectrogram similarity matching result is greater than or equal to 0.8, it is determined that the similarity of the spectrograms meets the detection conditions.

[0159] By analyzing the similarity matching results of the frequency domain graphs of the source audio file and the recorded audio file, the inclusion relationship between the two frequency domain graphs can be determined. The signal intensity in the frequency domain graph is transformed using the Fast Fourier Transform (FFT) method, resulting in a set of complex coefficients. The difference in signal intensity distribution corresponding to the same frequency is calculated, and the percentage of frequencies with positive differences relative to the total number of frequencies is statistically analyzed to determine the inclusion relationship between the two.

[0160] After calculating the difference distribution of the similarity of the frequency domain graph, the percentage between the number of positive differences and the total number can be obtained. In this application, the similarity threshold of the frequency domain graph is set to 80%. In this application, the similarity thresholds of the spectrogram and the frequency domain graph can be set to equal values. When the percentage is greater than 80%, it is determined that there is an inclusion relationship between the two, and the detection condition of the frequency domain graph is met. When the percentage is less than 80%, it is determined that there is no inclusion relationship between the two, and the detection condition of the frequency domain graph is not met.

[0161] When both the similarity of the spectrogram and the similarity of the frequency domain graph are greater than the similarity threshold, that is, when the similarity value of the spectrogram is greater than 0.8 and the similarity percentage of the frequency domain graph is greater than 80%, the detection result corresponding to the audio file is determined to be normal audio playback.

[0162] In step 3042, in response to the fact that the first similarity and the second similarity do not meet the detection conditions, the detection result corresponding to the second audio file is determined to be an audio playback abnormality.

[0163] For example, when the similarity between the spectrogram and the frequency domain graph does not simultaneously meet their respective thresholds, the detection result for the audio file is determined to be abnormal audio playback. In this case, the audio may not be playing properly or there may be a playback delay.

[0164] In some embodiments, after determining in step 3042 that the detection result corresponding to the audio file is an audio playback anomaly, the method further includes: adjusting the first playback parameters of the first audio file in response to the first similarity being less than the first similarity threshold, wherein the types of the first playback parameters include: playback speed, playback start time, playback duration, and playback end time; and adjusting the second playback parameters of the first audio file in response to the second similarity being less than the second similarity threshold, wherein the second playback parameters include: loudness and pitch.

[0165] For example, if abnormal audio playback is determined, the playback parameters of the source audio need to be adjusted. This can be done by adjusting the playback speed, start time, duration, and end time to restore normal playback. Simultaneously, since the playback does not meet the similarity thresholds of the spectrogram and frequency domain diagrams, the frequency in the playback time domain and the signal strength in the frequency domain of the source audio also need to be adjusted accordingly. In this case, the loudness and pitch of the source audio can be adjusted to restore normal playback.

[0166] In this embodiment, an audio file in the test environment is recorded, and the spectrogram and frequency domain diagram of the source audio file and the recorded audio file are obtained respectively. The similarity between the spectrogram and frequency domain diagram of the source audio file and the recorded audio file is compared. When the detection conditions of both the spectrogram and frequency domain diagram are met, it is determined that the audio file can be played correctly. Otherwise, the playback parameters are adjusted according to the comparison results of the spectrogram and frequency domain diagram similarity. This can improve the accuracy of audio detection and provide a basis for parameter adjustment when the audio is delayed or not played.

[0167] The following will describe an exemplary application of the audio detection method of this application in a real-world application scenario.

[0168] In related technologies, audio testing primarily relies on manual verification. The testing process involves human users experiencing the game and using their hearing to check the correctness and latency of the audio. This intuitive assessment determines whether the music plays correctly and whether playback begins immediately after the user clicks the play button, without any delay. Most non-human audio matching algorithms are applied to human voice matching, meaning they process primarily human voice features. However, in game scenarios or other application scenarios, the audio features are much richer, incorporating background music and various prompts. Commonly used human voice audio matching methods often fail to achieve the desired results, exhibiting low efficiency and accuracy in detecting the correctness of audio with multiple features.

[0169] There are many uncertainties in the process of testing related technologies using human senses. Because everyone's sensory experience is different and their hearing range is also different, it is impossible to rely entirely on human testing to accurately quantify the test results, nor can it obtain comparison results between the similarity and delay between the played audio and the original audio.

[0170] This application embodiment records audio in an application scenario, obtains the spectrogram and frequency domain diagram of the source audio file and the recorded audio file respectively, compares the similarity of the spectrogram and frequency domain diagram, quantifies the audio test, and combines the similarity matching results to obtain the results of audio correctness detection in multiple dimensions.

[0171] The following explanation is in conjunction with the accompanying drawings. Figure 4 , Figure 4 This is a schematic diagram of the sixth process of the audio detection method provided in this application embodiment. The executing entity can be a terminal device, a server, or a combination of both. This application embodiment takes a server as the executing entity as an example, and will combine... Figure 4 The steps shown are explained in detail.

[0172] In step 401, audio is recorded to obtain information for each audio call.

[0173] For example, in the process of automating testing of application scenarios, it is necessary to record audio throughout the entire testing process and collect audio information at the start and end of each audio playback to obtain information at each audio call.

[0174] The audio recording process includes using screen recording software to record the entire test process and noting the start and end times of the audio recording. During audio recording, playback relies on interfaces provided by the software or hardware environment that offers audio playback functionality. Therefore, the common library provides a general, reusable audio processing and playback function. This common library is an audio library written in JavaScript, a lightweight programming language commonly used in web development. The audio library provides rich functionality and interfaces for processing audio data, creating audio effects, and playing audio files, thus allowing for the delegation of the audio entry point class.

[0175] In the audio signal processing flow, a proxy class is used to manage and control the input of audio data. The proxy method provided by JavaScript is used to perform proxy operations. The object created by the proxy method can record the time and parameters of the method call each time the audio method is called, thereby obtaining information for each audio call.

[0176] For example, the recorded audio information includes: sound file path, sound duration (cost), sound playback start time (start_time), sound playback end time (end_time), and sound playback duration (duration). The sound playback duration attribute value is a floating-point number representing the total duration of the media file. The playback rate (playbackRate) represents the playback speed of the media file, allowing users to speed up or slow down the playback speed to achieve accelerated or slow-motion playback effects. The default value of the playback rate attribute is 1.0, representing normal playback speed. When the playback rate value is less than 1.0, the playback speed is slowed down; when the playback rate value is greater than 1.0, the playback speed is sped up. For example, a playback rate value of 0.5 means the audio is played at 0.5x speed; a playback rate value of 2.0 means the audio is played at 2.0x speed.

[0177] To facilitate understanding of the audio detection process described above, the overall audio detection workflow is introduced below. (See also...) Figure 5 , Figure 5 This is a schematic diagram illustrating the principle of the audio detection method provided in this application embodiment; the overall audio detection process can be achieved through... Figure 5 The steps 501 (audio recording), 502 (audio trimming), and 503 (audio matching) shown are implemented.

[0178] By recording audio in step 501, the recorded audio file (the second audio file mentioned above) can be obtained based on the source audio file. The audio of the entire test process can be recorded, and audio information such as audio file name, start time, duration and volume can be collected each time the audio is played and stopped. This provides a basis for audio trimming in step 502.

[0179] The entire running process of the game interface 510 is recorded to obtain the spectrum of the complete source audio file, represented as source audio file spectrum image 520. Based on the audio information, a recorded audio file is obtained, and the spectrum image 521 of the recorded audio file is obtained, representing a portion of the recorded source audio file spectrum within the 0-1.5 second time frame. After recording, all audio call information is compiled. For each audio call, audio trimming is performed in step 502 to obtain multiple trimmed sub-audio files. That is, the spectrum image 521 of the partially recorded source audio file is trimmed to obtain multiple independent trimmed sub-audio file spectrum images 522, which are used for audio matching in step 503.

[0180] The audio is verified through audio matching in step 503. For example, the similarity of the spectrograms of the cropped sub-audio file spectrum image 522 and the partially recorded source audio file spectrum image 521 is performed to verify whether the source audio is included in the recorded cropped audio.

[0181] In step 402, the recorded audio file is trimmed based on the audio information to obtain multiple sub-audio files.

[0182] For example, in the audio recording step, the start and end times of the audio are obtained, and each sub-audio segment is individually trimmed based on the start and end times to obtain multiple offline sub-audio files.

[0183] The audio trimming tool used in this application is FastForward Motion Picture Experts Group (FFMPEG), which can perform recording, conversion, encoding, and decoding operations. FFMPEG trims sub-audio files based on the start and end times of the audio playback. At the start of recording, time calibration is performed on all systems to ensure consistency. Therefore, the trimming is based on absolute time. Since absolute time is calculated based on the start time of the original audio file, it is unaffected by other factors (such as volume and sound effects).

[0184] In step 403, the spectrogram similarity between the audio source file and the sub-audio file is matched.

[0185] For example, after trimming based on audio information, multiple sub-audio files are obtained. A spectrogram of each sub-audio file with a time axis is plotted, as well as a spectrogram of the original audio file with a time axis. A spectrogram is a graphical representation of the signal's spectrum at various frequencies, using a wave-like pattern on the horizontal and vertical axes. It describes the signal's spectral information on a time-frequency plane, typically with time on the horizontal axis and frequency on the vertical axis, using different colors or grayscale values ​​to represent the intensity of different frequency components.

[0186] After obtaining the spectrograms of the sub-audio and the source audio file, a similarity matching process is performed between the two spectrograms. This matching process is based on an image similarity algorithm and uses a trained convolutional neural network. See also Figure 6 , Figure 6 This is a model structure diagram of the convolutional neural network provided in the embodiments of this application, which will be described in detail below.

[0187] During training, the convolutional neural network model 601 automatically learns the features of the images, extracts the feature vectors of the images for comparison, and evaluates the similarity of the images based on the comparison results. The normalization layer 6011 normalizes the input spectrogram images by scaling or cropping them to adjust the image size to a uniform resolution, ensuring that the images have the same dimensions. It also calculates the mean color value and normalizes the pixel values ​​of the images to the [0, 1] interval, ensuring that the two input spectrogram images have the same size and color space. The feature extraction layer 6012 performs convolution and pooling operations on the input images to extract feature vectors from the images, which can be pixel intensity, edges, texture, or shape. The feature comparison layer 6013 compares the feature vectors of the two images, using methods such as cosine similarity or Euclidean distance. The evaluation layer 6014 calculates the similarity between the two images based on the feature vector comparison results. The similarity is a value between 0 and 1, and the similarity value is used as the inclusion coefficient between the two images to represent their degree of similarity. The closer the similarity score is to 1, the more similar the two images are; the closer the similarity score is to 0, the less similar the two images are.

[0188] Based on the similarity matching results of the sub-audio and audio source file spectrograms, the inclusion relationship between the two spectrograms is obtained. The inclusion coefficient can be characterized by the similarity as described above. The inclusion relationship includes: complete inclusion, with an inclusion coefficient of 1, indicating that one spectrogram is completely contained in the other spectrogram, that is, one spectrogram covers all parts of the other spectrogram; partial overlap, with an inclusion coefficient close to 1, indicating that the two spectrograms partially overlap, that is, they have the same intensity at some frequencies; and no inclusion, with an inclusion coefficient of 0, indicating that one spectrogram has no overlap with the other spectrogram, that is, the frequency range and intensity distribution do not intersect.

[0189] In step 404, the frequency domain similarity of the audio source file and the sub-audio file is matched.

[0190] For example, in addition to matching the similarity of the spectrograms, statistical methods are also used to match the inclusion relationships of the frequency distribution curves, that is, to match the similarity of the pitch maps of the audio source file and the sub-audio files.

[0191] A frequency domain plot can represent the frequency distribution of a signal in the frequency domain. It is usually the result of a Fourier transform of the signal. The horizontal axis represents frequency and the vertical axis represents the amplitude or power of the signal. It is used to display the frequency components of the signal and can clearly show the frequency components of the signal. It only displays the frequency domain information of the signal and does not include time domain (time) information.

[0192] This application describes the Fast Fourier Transform (FFT) method for transforming audio signals, converting time-domain signals into frequency-domain signals. It is suitable for processing periodic signals, noise signals, and audio signals. The Fast Fourier Transform decomposes the signal sequence into frequency and phase information, enabling frequency domain analysis of the signal. The horizontal axis of the frequency distribution curve represents frequency, and the vertical axis represents the signal amplitude or power.

[0193] For example, to determine whether there is an inclusion relationship between the recorded audio and the source audio, when the ratio between the number of sampling points of the source audio frequency curve below the frequency curve of the recorded audio and the total number of sampling points of the source audio frequency curve is greater than a threshold (80%), it can be determined that there is an inclusion relationship between the two.

[0194] The specific calculation process is as follows: The input discrete signal sequence is transformed by a Fast Fourier Transform (FFT) to obtain a set of complex coefficients. These complex coefficients represent the signal components in the frequency domain, typically including real and imaginary parts, and can represent the amplitude and phase information of the signal at different frequencies. Each sampling point corresponds to a different frequency, starting from 0 Hz and reaching a maximum of 40,000 Hz (beyond 40,000 Hz, the human ear cannot distinguish the difference, so calculations are not performed). The difference between the recorded audio and the source audio distribution at each sampling point is calculated sequentially. For any given sampling point, if the difference is negative, it indicates that the source audio exceeds the recorded audio; if it is positive, it indicates that the recorded audio covers the source audio. Finally, the percentage between the number of sampling points where the source audio's frequency curve lies below the recorded audio's frequency curve and the total number of sampling points on the source audio's frequency curve is used as the inclusion factor in the frequency domain graph, thus determining the inclusion relationship between the two. When the percentage of the inclusion coefficient of the frequency domain plot is greater than 80%, it is determined that there is an inclusion relationship between the two; when the percentage of the inclusion coefficient of the frequency domain plot is less than 80%, it is determined that there is no inclusion relationship between the two.

[0195] See Figure 7 , Figure 7 This is a schematic diagram illustrating the inclusion relationship between the sub-audio and audio source file frequency domain diagrams provided in the embodiments of this application; Figure 7 The frequency domain diagram representing the recorded sub-audio 701 shows an inclusion relationship with the frequency domain diagram of the source audio 702. Region 7011 represents the frequency and signal intensity distribution of the recorded sub-audio 701, and region 7021 represents the signal intensity distribution corresponding to the frequencies of the source audio 702. The intensity distributions in regions 7021 and 7011 overlap, thus indicating a possible inclusion relationship. The inclusion coefficient ranges from 0 to 1; the closer it is to 1, the more likely the source audio is to be included in the recorded sub-audio. If the test results are normal, meaning the source audio is included in the recorded audio and the time axes of the audio are aligned, then the inclusion relationship is considered valid. Figure 7 The source audio 702 on the right is contained in the sub-audio 701 recorded on the left.

[0196] In step 405, the audio detection result is determined based on the inclusion relationship between the spectrogram and the frequency domain graph.

[0197] For example, after performing spectrogram similarity matching on the sub-audio and the audio source file, the inclusion relationship corresponding to the spectrogram is obtained. After performing frequency domain similarity matching on the sub-audio and the audio source file, the inclusion relationship corresponding to the frequency domain is obtained. When the inclusion relationship coefficient of the spectrogram is between 0.8 and 1, and the inclusion coefficient percentage of the frequency domain is greater than 80%, the audio detection result is determined to be that there is an inclusion relationship between the two audios. At this time, it indicates that the sound is played and the sound playback is not delayed.

[0198] The audio detection method provided in this application has the following beneficial effects:

[0199] By recording and trimming the audio during the testing process, multiple trimmed sub-audio files are obtained. The similarity between the sub-audio files and the original audio files is compared. By using both spectrogram and frequency domain similarity matching methods, matching results from multiple perspectives of audio information can be obtained. This helps to determine the relationship between the similarity and latency between audio files. Combining the inclusion relationship between audio files determined by the matching results of the two similarity methods can identify audio that has not been played and sound playback delay issues. This avoids the instability factors that rely on human experience, achieving fully automated audio correctness detection. It is a simple and efficient way to help small game projects discover audio-related problems.

[0200] The following description further illustrates the exemplary structure of the audio detection device 455 provided in this application embodiment as a software module. In some embodiments, as shown in FIG2, the software module stored in the audio detection device 455 in the memory 450 may include: an audio recording module 4551, used to acquire a first audio file and a second audio file to be detected, wherein the first audio file is a source audio file and the second audio file is recorded based on the first audio file; a similarity matching module 4552, used to determine a first similarity between the first audio file and the second audio file based on a first spectrogram of the first audio file and a second spectrogram of the second audio file; and to determine a second similarity between the first audio file and the second audio file based on a first frequency domain diagram of the first audio file and a second frequency domain diagram of the second audio file; and a detection result determination module 4553, used to determine the detection result corresponding to the second audio file based on the first similarity and the second similarity.

[0201] In some embodiments, the audio recording module 4551 is further configured to display a virtual scene in a human-computer interaction interface before acquiring the first audio file and the second audio file to be detected, wherein the virtual scene plays the first audio file when it runs; and in response to a trigger operation on the recording control, record the played first audio file to obtain the second audio file.

[0202] In some embodiments, the audio recording module 4551 is further configured to run the human-computer interaction interface via a mini-program embedded in the application, wherein the virtual scene is a mini-game virtual scene, and the recording is implemented through the audio output interface of the application; the display of the virtual scene in the human-computer interaction interface includes: in response to a trigger operation on the mini-program in the application, displaying the mini-game virtual scene associated with the mini-program.

[0203] In some embodiments, the audio recording module 4551 is further configured to record the first audio file being played when the virtual scene is running and the first audio file is in a playback state, to obtain a recording file, and to call the audio output interface to determine the playback start time, playback duration, and playback end time of the first audio file; and to trim the recording file based on the playback start time, the playback duration, and the playback end time to obtain at least one second audio file.

[0204] In some embodiments, the audio recording module 4551 is further configured to calibrate the following system times before recording the first audio file being played to obtain the second audio file: the first system time of the human-computer interaction interface, the second system time of the virtual scene, and the third system time of the audio output interface corresponding to the first audio file.

[0205] In some embodiments, the similarity matching module 4552 is further configured to determine a first spectrogram of the first audio file based on the frequencies of the first audio file at different times; determine a second spectrogram of the second audio file based on the frequencies of the second audio file at different times; and perform matching processing on the first spectrogram and the second spectrogram to obtain a first similarity between the first audio file and the second audio file.

[0206] In some embodiments, the similarity matching module 4552 is further configured to normalize the first spectrogram and the second spectrogram respectively to obtain a normalized first spectrogram and a normalized second spectrogram; perform feature extraction processing on the normalized first spectrogram to obtain a first image feature; perform feature extraction processing on the normalized second spectrogram to obtain a second image feature; determine a third similarity between the first image feature and the second image feature; and use the third similarity as a first similarity between the first audio file and the second audio file.

[0207] In some embodiments, the normalization process, the feature extraction process, and the process of determining the third similarity are implemented by a convolutional neural network model. The similarity matching module 4552 is further configured to, when the first image feature and the second image feature are represented as feature vectors, the third similarity is the cosine similarity between the first image feature and the second image feature.

[0208] In some embodiments, the similarity matching module 4552 is further configured to perform a Fourier transform on the audio time-domain signal of the first audio file to obtain a first frequency-domain signal; perform a Fourier transform on the audio time-domain signal of the second audio file to obtain a second frequency-domain signal; draw a first frequency-domain graph based on the first frequency-domain signal and draw a second frequency-domain graph based on the second frequency-domain signal, wherein the horizontal axis of the first frequency-domain graph and the second frequency-domain graph represents frequency and the vertical axis represents signal intensity; and determine a second similarity between the first audio file and the second audio file based on the difference in signal intensity distribution between the first frequency-domain graph and the second frequency-domain graph.

[0209] In some embodiments, the similarity matching module 4552 is further configured to perform the following processing for each frequency that is the same in the first frequency domain map and the second frequency domain map: determine the signal strength difference between the second signal strength of the frequency in the second frequency domain map and the first signal strength in the first frequency domain map; count a first number of positive signal strength differences; and take the ratio between the first number and the total number of signal strength differences as the second similarity between the first audio file and the second audio file.

[0210] In some embodiments, the detection result determination module 4553 is further configured to determine the detection result corresponding to the second audio file as normal audio playback in response to the first similarity and the second similarity satisfying the detection conditions, wherein the detection conditions include: the first similarity is greater than a first similarity threshold and the second similarity is greater than a second similarity threshold; and to determine the detection result corresponding to the second audio file as abnormal audio playback in response to the first similarity and the second similarity not satisfying the detection conditions.

[0211] In some embodiments, the detection result determination module 4553 is further configured to, after determining that the detection result corresponding to the second audio file is an audio playback abnormality, adjust the first playback parameters of the first audio file in response to the first similarity being less than the first similarity threshold, wherein the types of the first playback parameters include: playback speed, playback start time, playback duration, and playback end time; and adjust the second playback parameters of the first audio file in response to the second similarity being less than the second similarity threshold, wherein the second playback parameters include: loudness and pitch.

[0212] This application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the audio detection method described above in this application.

[0213] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the audio detection method provided in this application. For example, ... Figure 3A The audio detection method is shown.

[0214] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.

[0215] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.

[0216] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).

[0217] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.

[0218] In summary, by comparing the similarity of the source audio and the recorded audio file in both spectrogram and frequency domain diagrams through the embodiments of this application, and simultaneously satisfying the detection conditions of both similarities, it is determined that the audio file can be played normally, providing a basis for adjusting audio playback parameters and improving the accuracy of detecting whether the audio is played normally and whether there is a delay.

[0219] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. An audio detection method, characterized in that, The method comprises: Obtain a first audio file and a second audio file to be detected, wherein the first audio file is the source audio file and the second audio file is recorded based on the first audio file; Based on the first spectrogram of the first audio file and the second spectrogram of the second audio file, a first similarity between the first audio file and the second audio file is determined. Based on the first frequency domain diagram of the first audio file and the second frequency domain diagram of the second audio file, a second similarity between the first audio file and the second audio file is determined. Based on the first similarity and the second similarity, the detection result corresponding to the second audio file is determined.

2. The method according to claim 1, characterized in that, Before acquiring the first audio file and the second audio file to be detected, the method further includes: A virtual scene is displayed in the human-computer interaction interface, wherein the virtual scene plays a first audio file when it runs; In response to a trigger operation on the recording control, the first audio file being played is recorded to obtain a second audio file.

3. The method according to claim 2, characterized in that, The human-computer interaction interface runs through a mini-program embedded in the application, the virtual scene is a mini-game virtual scene, and the recording is achieved through the audio output interface of the application; The process of displaying a virtual scene in the human-computer interaction interface includes: In response to a triggered operation on a mini-program within the application, a virtual scene of a mini-game associated with the mini-program is displayed.

4. The method according to claim 2, characterized in that, The process of recording the first audio file to obtain the second audio file includes: When the virtual scene is running and the first audio file is playing, the playing first audio file is recorded to obtain a recording file, and Call the audio output interface to determine the start time, duration, and end time of playback for the first audio file; Based on the playback start time, the playback duration, and the playback end time, the recorded file is trimmed to obtain at least one second audio file.

5. The method according to claim 2, characterized in that, Before recording the first audio file to obtain the second audio file, the method further includes: The following system times are calibrated: the first system time of the human-computer interaction interface, the second system time of the virtual scene, and the third system time of the audio output interface corresponding to the first audio file.

6. The method according to claim 1, characterized in that, Determining the first similarity between the first audio file and the second audio file based on the first spectrogram of the first audio file and the second spectrogram of the second audio file includes: Based on the frequencies of the first audio file at different times, a first spectrogram of the first audio file is determined. Based on the frequencies of the second audio file at different times, a second spectrogram of the second audio file is determined; The first spectrogram and the second spectrogram are matched to obtain a first similarity between the first audio file and the second audio file.

7. The method according to claim 6, characterized in that, The matching process of the first spectrogram and the second spectrogram to obtain the first similarity between the first audio file and the second audio file includes: The first and second spectrograms are normalized respectively to obtain the normalized first spectrogram and the normalized second spectrogram; The first normalized spectrogram is subjected to feature extraction processing to obtain the first image features; The normalized second spectrogram is subjected to feature extraction processing to obtain the second image features; Determine a third similarity between the first image feature and the second image feature; The third similarity is used as the first similarity between the first audio file and the second audio file.

8. The method according to claim 7, characterized in that, The normalization process, the feature extraction process, and the process for determining the third similarity are implemented through a convolutional neural network model. When the first image feature and the second image feature are represented as feature vectors, the third similarity is the cosine similarity between the first image feature and the second image feature.

9. The method according to claim 1, characterized in that, Determining the second similarity between the first audio file and the second audio file based on the first frequency domain map of the first audio file and the second frequency domain map of the second audio file includes: Perform a Fourier transform on the audio time-domain signal of the first audio file to obtain the first frequency-domain signal; Perform a Fourier transform on the audio time-domain signal of the second audio file to obtain the second frequency-domain signal; A first frequency domain diagram is drawn based on the first frequency domain signal, and a second frequency domain diagram is drawn based on the second frequency domain signal, wherein the horizontal axis of the first frequency domain diagram and the second frequency domain diagram represents frequency, and the vertical axis represents signal strength; Based on the difference in signal intensity distribution between the first frequency domain map and the second frequency domain map, a second similarity is determined between the first audio file and the second audio file.

10. The method according to claim 9, characterized in that, Determining the second similarity between the first audio file and the second audio file based on the difference in signal intensity distribution between the first frequency domain map and the second frequency domain map includes: For each frequency that is the same in the first frequency domain diagram and the second frequency domain diagram, the following process is performed: determine the signal strength difference between the second signal strength of the frequency in the second frequency domain diagram and the first signal strength in the first frequency domain diagram; Count the first number of positive signal strength differences; The ratio between the first quantity and the total number of signal strength differences is used as the second similarity between the first audio file and the second audio file.

11. The method according to any one of claims 1 to 10, characterized in that, The step of determining the detection result corresponding to the second audio file based on the first similarity and the second similarity includes: In response to the first similarity and the second similarity satisfying the detection conditions, the detection result corresponding to the second audio file is determined to be normal audio playback, wherein the detection conditions include: the first similarity is greater than the first similarity threshold, and the second similarity is greater than the second similarity threshold; In response to the fact that the first similarity and the second similarity do not meet the detection conditions, the detection result corresponding to the second audio file is determined to be an audio playback abnormality.

12. The method according to claim 11, characterized in that, After determining that the detection result corresponding to the second audio file is an audio playback abnormality, the method further includes: In response to the first similarity being less than the first similarity threshold, the first playback parameters of the first audio file are adjusted. The types of the first playback parameters include: playback speed, playback start time, playback duration, and playback end time. In response to the second similarity being less than the second similarity threshold, the second playback parameters of the first audio file are adjusted, the second playback parameters including loudness and pitch.

13. An audio detection device, characterized in that, The device comprises: An audio acquisition module is used to acquire a first audio file and a second audio file to be detected, wherein the first audio file is a source audio file and the second audio file is recorded based on the first audio file; The similarity matching module is used to determine a first similarity between the first audio file and the second audio file based on a first spectrogram of the first audio file and a second spectrogram of the second audio file; and to determine a second similarity between the first audio file and the second audio file based on a first frequency domain diagram of the first audio file and a second frequency domain diagram of the second audio file. The detection result determination module is used to determine the detection result corresponding to the second audio file based on the first similarity and the second similarity.

14. An electronic device, characterized in that, The electronic device includes: Memory is used to store executable instructions for a computer; A processor, when executing computer-executable instructions stored in the memory, implements the audio detection method according to any one of claims 1 to 12.

15. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, the audio detection method according to any one of claims 1 to 12 is implemented.