Audio recognition method, device, equipment, medium and program product
By periodically sending audio data and converting encoding formats in a weak network environment, from PCM to MP3, the problem of low audio recognition efficiency in a weak network environment is solved, and efficient audio recognition and timely acquisition of recognition results are achieved.
Patent Information
- Application Number
- CN202510558015.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-08
AI Technical Summary
In a weak network environment, the existing audio recognition technology has high transmission delay, data loss due to the large amount of PCM data, and is unable to receive server feedback results in a timely manner.
By periodically sending audio data, according to the server's reply feedback, the encoding format is converted from PCM to MP3 in a weak network environment, reducing the amount of data and improving transmission efficiency.
Improve the success rate and efficiency of audio recognition in a weak network environment, avoid information transmission timeout, and ensure that the client receives recognition results in a timely manner.
Smart Images

Figure CN120279943A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of computer technology, and in particular, to an audio recognition method, apparatus, device, medium, and program product. Background Art
[0002] Song identification by listening is a technology that uses audio processing and pattern recognition technologies to analyze an audio clip input by a user, so as to identify information such as the name of the corresponding song and the singer. Most music playback applications or voice assistant applications can provide the song identification by listening function, providing users with a convenient music retrieval service.
[0003] In the related art, the client collects an audio signal through the microphone of the device and converts it into Pulse Code Modulation (PCM) data. PCM data is an uncompressed digital audio format that discretizes the audio signal at a fixed sampling rate and bit depth, completely retaining the waveform information of the original audio. After the client sends the collected PCM data to the server for feature analysis and compares it with the song features in the database, the matching result is finally returned to the client.
[0004] However, the above method has limitations in a weak network environment, which is usually a current network environment with high network latency and low network transmission rate, such as scenarios in remote areas and underground parking lots. Due to the large amount of PCM data, data loss, information transmission timeout and other problems are likely to occur during transmission in a weak network environment. The client cannot receive the result returned by the server in time, resulting in song identification failure and low song identification efficiency. Summary of the Invention
[0005] Embodiments of the present application provide an audio recognition method, apparatus, device, medium, and program product, which can improve the audio recognition efficiency and the rate of audio data transmission. The technical solution is as follows:
[0006] On the one hand, an audio recognition method is provided, and the method includes:
[0007] Obtain first audio data, where the first audio data is the data to be subjected to audio recognition;
[0008] Send the first audio data to the server in a first encoding format according to a preset sending period, where the first audio data is used to instruct the server to perform audio recognition;
[0009] In the case that the reply feedback of the first audio data sent by the server for the first n preset sending cycles meets the weak network requirements, starting from the (n + 1)-th preset sending cycle, the first audio data is sent to the server in a second coding format, where, when the first audio data is the same, the amount of data of the coding data result corresponding to the second coding format is less than the amount of data of the coding data result corresponding to the first coding format, and n is a positive integer.
[0010] On the other hand, an audio recognition device is provided, and the device includes:
[0011] An acquisition module, configured to acquire first audio data, where the first audio data is data to be subjected to audio recognition;
[0012] A sending module, configured to send the first audio data to the server in a first coding format according to a preset sending cycle, where the first audio data is used to instruct the server to perform audio recognition;
[0013] The sending module is further configured to, in the case that the reply feedback of the first audio data sent by the server for the first n preset sending cycles meets the weak network requirements, starting from the (n + 1)-th preset sending cycle, send the first audio data to the server in a second coding format, where, when the first audio data is the same, the amount of data of the coding data result corresponding to the second coding format is less than the amount of data of the coding data result corresponding to the first coding format, and n is a positive integer.
[0014] In an optional embodiment, the sending module is further configured to, in the case that the number of reply times of the reply feedback of the first audio data sent by the server for the first n preset sending cycles does not reach a first number threshold, determine that the reply feedback meets the weak network requirements; and starting from the (n + 1)-th preset sending cycle, send the first audio data to the server in the second coding format.
[0015] In an optional embodiment, the sending module is further configured to send the first audio data to the server in the first coding format according to the preset sending cycle through a first communication interface;
[0016] The sending module is further configured to, starting from the (n + 1)-th preset sending cycle, send the first audio data to the server in the second coding format through a second communication interface; where the first communication interface and the second communication interface are different interfaces for data interaction between the client and the server.
[0017] In an optional embodiment, the sending module is further configured to start from the (n + 1)-th preset sending period, perform format conversion on the first audio data, change the format of the first audio data from the first encoding format to the second encoding format; and send the first audio data to the server in the second encoding format.
[0018] In an optional embodiment, the sending module is further configured to perform at least one compression process on the first audio data to obtain at least one compressed audio data; obtain the compression ratio of each of the at least one compressed audio data with respect to the first audio data; stop the compression process in response to the compression ratio corresponding to the currently obtained compressed audio data meeting a preset format conversion requirement; and determine the currently obtained compressed audio data as the first audio data corresponding to the second encoding format.
[0019] In an optional embodiment, the sending module is further configured to perform the i-th compression process on the first audio data to obtain the i-th compressed audio data; calculate the data volume difference between the i-th compressed audio data and the first audio data to obtain the i-th difference audio data, where i is a positive integer; and obtain the i-th compression ratio based on the ratio between the i-th difference audio data and the first audio data.
[0020] In an optional embodiment, the sending module is further configured to stop the compression process when the (k - 1)-th compression ratio obtained after the (k - 1)-th compression process does not reach a preset compression ratio threshold and the k-th compression ratio obtained after the k-th compression process reaches the preset compression ratio threshold; where k is a positive integer; and determine the compressed audio data corresponding to the k-th compression ratio as the first audio data corresponding to the second encoding format.
[0021] In an optional embodiment, the obtaining module is further configured to collect audio data segments based on the preset sending period to obtain the first audio data, where the duration of the audio data segments matches the period duration of the preset sending period.
[0022] The sending module is further configured to, in the n-th preset sending period, integrate the audio data segments collected respectively in the previous n preset sending periods to obtain the n-th integrated audio data, and use it as the first audio data sent in the n-th preset sending period; and send the n-th integrated audio data to the server in the first encoding format in the n-th preset sending period.
[0023] In an optional embodiment, the obtaining module is further configured to obtain signal strength information and location information of the terminal corresponding to the client when obtaining the first audio data, where the signal strength information is used to indicate the signal strength of the environment where the terminal is located; determine network status information based on the location information and the signal strength information, where the network status information is used to indicate the network status when the client transmits data;
[0024] The sending module is further configured to, when the network status information in the previous n preset sending cycles meets the weak network requirement, start from the (n + 1)-th preset sending cycle, and send the first audio data to the server in the second coding format.
[0025] On the other hand, a computer device is provided, where the computer device includes a processor and a memory, and at least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the audio recognition method according to any one of the above embodiments of the present application.
[0026] On the other hand, a computer-readable storage medium is provided, and at least one instruction, at least one program, a code set or an instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the audio recognition method according to any one of the above embodiments of the present application.
[0027] On the other hand, a computer program product or a computer program is provided, where the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the audio recognition method according to any one of the above embodiments.
[0028] The beneficial effects brought by the technical solutions provided in the embodiments of the present application at least include:
[0029] In the process of audio recognition, by periodically sending the audio data to be recognized, it is determined whether the current network environment is a weak network environment with a low transmission rate according to the reply feedback situation of the server within a specified number of cycles. The client determines the encoding format used when sending audio data to the server according to the network environment. It can convert the transmission format from the first encoding format to the second encoding format with a smaller data volume in the case of a weak network environment, reduce the transmission data volume of the audio data, reduce the amount of information requests from the client, and thus improve the transmission efficiency of the audio data. It avoids the situation of information transmission timeout caused by poor weak network environment and the client being unable to receive the server feedback result in time, and improves the success rate of audio recognition and the efficiency of obtaining audio recognition results. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following described drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0031] Figure 1 is a schematic diagram of an audio recognition system provided by an exemplary embodiment of the present application;
[0032] Figure 2 is a flowchart of an audio recognition method provided by an exemplary embodiment of the present application;
[0033] Figure 3 is a flowchart of an audio recognition method provided by another exemplary embodiment of the present application;
[0034] Figure 4 is a flowchart of a song recognition method provided by an exemplary embodiment of the present application;
[0035] Figure 5 is a schematic diagram of the song recognition interface of the client provided by an exemplary embodiment of the present application;
[0036] Figure 6 is a block diagram of the structure of an audio recognition device provided by an exemplary embodiment of the present application;
[0037] Figure 7 is a block diagram of the structure of a computer device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the drawings.
[0039] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms "a", "the", and "said" used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0040] It should be noted that the information and data involved in this application are all information and data authorized by users or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0041] First, a brief introduction to the nouns involved in the embodiments of this application:
[0042] Pulse Code Modulation (PCM) audio: It is a coding method that converts analog audio signals into digital signals. By sampling, quantizing, and coding the audio signals, they are converted into a series of digital pulse signals for storage, processing, and transmission by a computer or other digital devices.
[0043] In this application, the first coding format refers to the PCM format, which is an uncompressed audio format.
[0044] Moving Picture Experts Group Audio Layer 3 (MP3) audio: A lossy compressed digital audio format that can compress the original audio to 1 / 10 of its size while retaining the main sound quality perceptible to the human ear.
[0045] In this application, the second coding format refers to the MP3 format. The first audio data sent in the first coding format is PCM audio data, and the first audio data sent in the second coding format is MP3 audio data. That is, the MP3 audio data is obtained by compressing the PCM audio data, realizing the conversion from the first coding format to the second coding format.
[0046] Weak network: Generally refers to the current network environment with high network latency and low network transmission rate.
[0047] Song recognition by listening: It is a function that recognizes music by transmitting a PCM audio byte array to the server. It allows users to search for and identify information such as the name of the song, the singer, and the lyrics by recording or playing a music clip.
[0048] Secondly, an explanation of the audio recognition system involved in the embodiments of this application. Schematically, please refer to Figure 1, the audio recognition system involves a client 110 and a server 120. The client 110 and the server 120 are connected through a communication network 100, and there are two different data interaction interfaces between the client 110 and the server 120: a first communication interface and a second communication interface, which are respectively used to transmit audio data in different encoding formats.
[0049] Taking the process of identifying a song by listening as an example to illustrate the audio recognition system of the present application. Among them, the client 110 can provide the function of identifying a song by listening.
[0050] A control for triggering the function of identifying a song by listening is displayed on the display interface of the client 110. After receiving the trigger operation of the user for the control, the client 110 starts to obtain the first audio data. Among them, the first audio data is the data to be recognized for audio.
[0051] During the process of obtaining the first audio data, the client 110 will first send the first audio data in the first encoding format to the server 120 based on a preset sending period. After the server 120 recognizes the first audio data in the first encoding format, it returns the recognition result to the client 110.
[0052] Within the first n preset cycles, the client 110 will send the first audio data n times in the first encoding format through the first communication interface, and the duration of the first audio data sent each time gradually increases. For example, if the cycle duration is 3 seconds and the entire process of identifying a song by listening lasts for 15 seconds, then within 15 seconds, the client 110 will continuously collect the first audio data.
[0053] The first cycle is from 0 to 3 seconds. The first audio data is sent for the first time at the 3rd second, with a duration of 3 seconds; the second cycle is from 3 to 6 seconds. The first audio data is sent for the second time at the 6th second, with a duration of 6 seconds; and so on.
[0054] That is to say, each time the first audio data sent is the audio data obtained by taking the moment when the first audio data starts to be collected as the starting timestamp and the moment when the data is sent as the ending timestamp.
[0055] In the (n + 1)th preset cycle, the client 110 will determine whether to change the format of sending the first audio data according to the return result of the server 120.
[0056] If the number of return results received within the first n preset cycles does not reach n times, it means that the current network environment is a weak network environment, which meets the weak network requirements, and it is necessary to convert the format of the first audio data. Starting from the (n + 1)th preset cycle, the first audio data is sent in the second encoding format.
[0057] Exemplarily, the client 110 maintains a weak network identifier, and the value of the weak network identifier is used to indicate whether the current network environment is a weak network environment.
[0058] The client 110 determines whether the value of the weak network identifier is True or False according to the reply feedback situation of the server 120 for the previous n preset transmission cycles (that is, the number of times the return result is received by the client 110).
[0059] When the value of the weak network identifier is True, it means that the preset weak network requirements are met. Starting from the (n + 1)-th preset cycle, the client 110 sends the first audio data to the server 120 through the second communication interface in the second encoding format. After the server 120 receives the first audio data with an increased duration based on the second communication interface, it performs identification and analysis based on the newly received first audio data and returns the identification result.
[0060] When the value of the weak network identifier is False, it means that the preset weak network requirements are not met. Then the client 110 still uses the method of sending the first audio data within the previous n preset cycles to send the first audio data in the first encoding format through the first communication interface.
[0061] After the server 120 receives the first audio data with an increased duration based on the first communication interface, it performs identification and analysis based on the newly received first audio data and returns the identification result.
[0062] Among them, the process of the server 120 performing identification processing on the first audio data is mainly as follows: First, it receives the song recognition request sent by the client 110 to obtain the first audio data. First, it judges whether the encoding format of the first audio data is the second encoding format.
[0063] That is, the server 120 judges whether the interface used to receive the first audio data is the second communication interface.
[0064] If it is the second communication interface, it means that the format of the first audio data is the second encoding format, and the first audio data needs to be decompressed to obtain the first audio data in the first encoding format, and then identification and analysis are performed.
[0065] If it is the first communication interface, it means that the format of the first audio data is the first encoding format, and the first audio data can be directly identified and analyzed.
[0066] For the first audio data in the first encoding format, the server 120 extracts the audio fingerprint features of the first audio data. The audio fingerprint features refer to a set of key mathematical features extracted from the audio, which can uniquely identify this audio.
[0067] The server 120 stores an audio database, which contains multiple song audios and corresponding audio information. After the server 120 extracts the fingerprint features of the multiple song audios in the audio database in advance, the audio fingerprint features corresponding to each song are obtained and stored.
[0068] The audio fingerprint features of the first audio data are used as retrieval conditions to match the audio fingerprint features of multiple songs pre-stored in the server 120, and the song audio with the highest matching degree is recalled from the audio database as the return result.
[0069] The return result contains the song audio and audio information, which is used to indicate that the first audio data is the data collected from the same song audio.
[0070] After receiving the return result, the client 110 will display the information of the song on the display interface, such as the song name and the name of the singer who sings the song.
[0071] In some embodiments, if the first audio data is from an audio that is being played and still playing, the client 110 will align the current playing progress according to the audio information returned by the server 120 and play the song synchronously.
[0072] The device running the client 110 described above can be various forms of terminal devices such as mobile phones, tablet computers, desktop computers, portable laptops, smart TVs, vehicle-mounted terminals, and smart home devices, and the embodiments of the present application do not limit this.
[0073] It should be noted that the above server 120 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0074] In some embodiments, the above server 120 can also be implemented as a node in a blockchain system.
[0075] Combined with the above noun introduction and application scenarios, the audio recognition method provided by the present application is described. This method can be executed by the server or the client, or jointly executed by the server and the client. In the embodiments of the present application, this method is described by taking the execution by the client as an example, as Figure 2 shown, Figure 2 is the flowchart of the audio recognition method provided by an exemplary embodiment of the present application. This method includes the following steps.
[0076] Step 210, obtain the first audio data.
[0077] Among them, the client can provide audio recognition functions, and the types of clients include, but are not limited to, the following:
[0078] (1) Mobile applications: For example, application programs / software installed on devices such as mobile phones, smart watches / bracelets, etc.;
[0079] (2) Desktop applications: For example, application programs / software running on computers or browser web pages;
[0080] (3) Smart hardware devices: For example, smart speakers, voice recorders, in-vehicle systems, etc.
[0081] Among them, the first audio data is the data to be recognized by audio.
[0082] The types of the first audio data include, but are not limited to, the following:
[0083] (1) Audio collected in real time by the client. For example, audio collected by the client through the microphone or other audio collection components of the terminal at the current moment;
[0084] (2) Audio cached locally on the terminal to which the client belongs. This audio has a certain duration and the user can customize the playback progress of the audio.
[0085] Step 220, send the first audio data to the server in the first encoding format according to a preset sending period.
[0086] Sending the first audio data according to a preset sending period means that, based on a specified duration as the cycle duration, the first audio data is sent to the server side at intervals of the cycle duration, and the time interval between adjacent two sending steps is the cycle duration.
[0087] Among them, the first audio data in the first encoding format is audio data in the Pulse Code Modulation (PCM) format. That is to say, sending the first audio data to the server in the first encoding format actually means sending PCM audio data to the server.
[0088] Among them, the PCM audio data is represented as a continuous byte array, in which the quantized amplitude values of each sampling point are stored. That is to say, each value stored in the array refers to the amplitude value of each sampling point, representing the instantaneous amplitude of the sound wave at a certain moment.
[0089] A sampling point refers to a time node at which the continuous analog audio signal is sampled at a fixed time interval during the process of converting it into a discrete digital signal.
[0090] For example, in a piece of audio, when the sound wave vibrates strongly at a certain moment, the amplitude value of the corresponding sampling point is large; when the vibration is weak, the value is small.
[0091] Exemplarily, taking PCM audio data with a 16-bit quantization precision, mono-channel, and a sampling frequency of 8 kHz as an example, the amplitude value of each sampling point occupies 2 bytes (16 bits).
[0092] Suppose a certain PCM byte array is [0x01, 0x00, 0x03, 0xFE, 0x05, 0x00,...], where every two bytes form a group, representing the quantization amplitude value of a sampling point. For example, 0x0100 (hexadecimal) is converted to decimal as 256, and 0x03FE is converted to decimal as 1022. These values represent the quantization results of the sound wave vibration amplitudes at different sampling points. The array is arranged in chronological order, completely recording the waveform change characteristics of the first audio data.
[0093] Among them, the first audio data is used to instruct the server to perform audio recognition.
[0094] Exemplarily, taking the scenario of identifying a song by listening as an example, the audio recognition process executed by the server after receiving the first audio data is to determine the relevant information of the song corresponding to the first audio data, including but not limited to: song name, singer name, lyric information, etc.
[0095] Optionally, the first audio data is sent to the server through the first communication interface in the first encoding format according to a preset sending period.
[0096] The first communication interface is an interface for data interaction between the client and the server. When the client transmits audio data in the first encoding format, it will call the first communication interface. The server receives the first audio data through the first communication interface. At this time, the encoding format of the first audio data can be directly determined as the PCM format according to the type of the first communication interface.
[0097] Optionally, the first audio data is audio data collected in real time, and the previously collected first audio data is sent to the server in real time at intervals of each cycle duration.
[0098] Based on a preset sending period, audio data segments are collected to obtain the first audio data.
[0099] Among them, the duration of the audio data segment matches the cycle duration of the preset sending period, that is, the duration of the audio data segment collected in the a-th preset sending period is equal to the duration of the a-th preset sending period. a is a positive integer.
[0100] Exemplarily, in the nth preset transmission period, the audio data segments collected in the previous n preset transmission periods are integrated to obtain the nth integrated audio data, which is used as the first audio data transmitted in the nth preset transmission period.
[0101] In the nth preset transmission period, the nth integrated audio data is sent to the server in the first encoding format.
[0102] For example, if the period duration of the preset transmission period is 3 seconds, the first audio data is sent to the server every 3 seconds, and the duration of each audio data segment is also 3 seconds.
[0103] The entire audio recognition process lasts for 15 seconds. Within 15 seconds, the client 110 continuously collects the first audio data, and the total duration of the finally collected first audio data is 15 seconds.
[0104] The first cycle is from 0 to 3 seconds. The first audio data is sent for the first time at the 3rd second, with a duration of 3 seconds. The second cycle is from 3 to 6 seconds. The first audio data is sent for the second time at the 6th second, with a duration of 6 seconds. And so on. That is, each time the first audio data sent is the audio data obtained with the start time of collecting the first audio data as the start timestamp and the time of sending the data as the end timestamp.
[0105] Step 230, when the reply feedback from the server for the first audio data sent in the previous n preset transmission periods meets the weak network requirements, starting from the (n + 1)th preset transmission period, the first audio data is sent to the server in the second encoding format.
[0106] Among them, when the first audio data is the same, the amount of data of the encoding data result corresponding to the second encoding format is less than the amount of data of the encoding data result corresponding to the first encoding format, and n is a positive integer.
[0107] Exemplarily, the first audio data in the second encoding format refers to the audio data in the Moving Picture Experts Group Audio Layer III (MP3) format. That is, sending the first audio data to the server in the second encoding format actually means sending MP3 audio data to the server.
[0108] Optionally, when the number of reply times of the reply feedback from the server for the first audio data sent in the previous n preset transmission periods does not reach the first number threshold, it is determined that the reply feedback meets the weak network requirements.
[0109] Starting from the (n + 1)th preset transmission period, the first audio data is sent to the server in the second encoding format.
[0110] The reply feedback meeting the weak network requirement means that the current transmission network environment between the client and the server is a weak network environment, with high network latency and low network transmission rate, resulting in the server being unable to feedback the recognition result to the client within the specified time, and the client not completely receiving the reply feedback from the server.
[0111] Each time the client sends the first audio data, it means that the client initiates an information request, requesting the server to perform audio recognition based on the received first audio data.
[0112] For each information request initiated by the client, the server will return a recognition result as a reply feedback each time.
[0113] The number of times the client receives the recognition result is regarded as the number of replies for the server to reply to the client's request.
[0114] Exemplarily, the first number threshold is n.
[0115] That is, within the first n preset sending cycles, the client sends the first audio data n times. If the client receives the reply feedback from the server n times, it means that the number of replies from the server reaches the first number threshold and does not meet the weak network requirement.
[0116] If the client receives the reply feedback from the server m times (m is less than n), it means that the number of replies from the server does not reach the first number threshold and meets the weak network requirement.
[0117] For example, n is 2, and the duration of the preset sending cycle is 3 seconds. Then within the first 2 preset sending cycles, the client sends the first audio data in the first coding format through the first communication interface at the 3rd second and the 6th second respectively. Before sending the first audio data at the 9th second, first obtain the reply feedback situation of the server in the past 6 seconds.
[0118] If the number of replies reaches 2 times, it means that it does not meet the weak network requirement, and still sends the first audio data in the first coding format to the server through the first communication interface at the 9th second.
[0119] If the number of replies does not reach 2 times (1 time or 0 time), it means that it meets the weak network requirement, and sends the first audio data in the second coding format to the server through the second communication interface at the 9th second.
[0120] Optionally, starting from the (n + 1)th preset sending cycle, send the first audio data to the server through the second communication interface in the second coding format.
[0121] Among them, the first communication interface and the second communication interface are different interfaces for data interaction between the client and the server.
[0122] That is, within the first n preset transmission cycles, the purpose of audio data transmission between the client and the server is to determine whether the current network environment meets the requirements of a weak network, and then determine the communication interface type and encoding format used by the client to send the first audio data to the server starting from the (n + 1)-th preset transmission cycle.
[0123] Optionally, when the reply feedback from the server for the first audio data sent during the first n preset transmission cycles meets the requirements of a weak network, starting from the (n + 1)-th preset transmission cycle, format conversion is performed on the first audio data, and the format of the first audio data is changed from the first encoding format to the second encoding format. The first audio data is sent to the server in the second encoding format.
[0124] Exemplarily, at least one compression process is performed on the first audio data to obtain at least one compressed audio data.
[0125] Among them, the compressed audio data refers to the first audio data in the second encoding format, that is, MP3 audio data.
[0126] The ways of at least one compression process include the following two:
[0127] (1) Each compression is performed on the original first audio data;
[0128] For example, during the first compression process, the first audio data is compressed to obtain the first compressed audio data; during the second compression process, the first audio data is still compressed to obtain the second compressed audio data, and so on. Among them, the data volume sizes corresponding to each compressed audio data are different.
[0129] (2) Each compression is performed by repeatedly compressing the audio data obtained from the previous compression process;
[0130] For example, during the first compression process, the first audio data is compressed to obtain the first compressed audio data; during the second compression process, based on the first compression process, the first compressed audio data is compressed again to obtain the second compressed audio data, and so on. Among them, the data volume sizes corresponding to each compressed audio data are different.
[0131] Optionally, the data volume size of the audio data obtained from the (t + 1)-th compression process is smaller than the data volume size of the audio data obtained from the t-th compression process, where t is a positive integer.
[0132] Obtain the compression ratios of at least one compressed audio data with respect to the first audio data respectively.
[0133] The compression ratio refers to the proportion of the data reduction amount of the compressed audio data compared to the first audio data in the first audio data. The higher the compression ratio, the smaller the data volume of the compressed audio data obtained after compressing the first audio data.
[0134] Optionally, when the compression ratio corresponding to the current compressed audio data obtained by the current compression process meets the preset format conversion requirements, stop the compression process.
[0135] Determine the current compressed audio data as the first audio data corresponding to the second encoding format.
[0136] Exemplarily, the preset format conversion requirement means that the compression ratio of the current audio data obtained after compression is higher than the preset compression ratio threshold. For example, the compression ratio threshold is 50%.
[0137] Exemplarily, perform the i-th compression process on the first audio data to obtain the i-th compressed audio data.
[0138] Calculate the data volume difference between the i-th compressed audio data and the first audio data to obtain the i-th difference audio data, where i is a positive integer.
[0139] Based on the ratio between the i-th difference audio data and the first audio data, obtain the i-th compression ratio.
[0140] Schematically, refer to Table 1 below. Table 1 is a comparison table before and after compression between a first audio data and compressed audio data, which includes the compression ratio of each compressed audio data relative to the first audio data.
[0141] Table 1
[0142]
[0143] Among them, Table 1 lists the data volume sizes before and after compressing the first audio data of different durations and the corresponding compression ratios.
[0144] For example, when the duration of the first audio data is 3 seconds, the data volume size before compression is 47k, and the data volume size of the compressed audio data obtained after compression is 11k. Then, it is calculated that (47 - 11) / 47 = 76.60%, which is the compression ratio, and so on.
[0145] Exemplarily, if the compression ratio threshold is 50%, the compression ratios of the compressed audio data obtained by compressing the first audio data of the above different durations all meet the requirements, and the first audio data can be sent to the server in the second encoding format through the second communication interface, that is, the compressed audio data is sent to the server through the second communication interface.
[0146] Optionally, when the (k - 1)th compression ratio obtained after the (k - 1)th compression process does not reach the preset compression ratio threshold and the kth compression ratio obtained after the kth compression process reaches the preset compression ratio threshold, the compression process is stopped. k is a positive integer.
[0147] Determine the compressed audio data corresponding to the kth compression ratio as the first audio data corresponding to the second encoding format.
[0148] Exemplarily, the preset compression ratio threshold is 50%.
[0149] Then, calculate the compression ratio of the compressed audio data corresponding to each compression process respectively. If the compressed audio data with a compression ratio reaching the preset compression ratio threshold appears for the first time, stop the compression process, use this compressed audio data as the first audio data of the second encoding format, and send it to the server.
[0150] For example, after the first audio data undergoes the first compression process, the compression ratio of the corresponding first compressed audio data is 30%, which does not reach the preset compression ratio threshold of 50%, so continue to compress.
[0151] After the first audio data undergoes the second compression process, the compression ratio of the corresponding second compressed audio data is 45%, which does not reach the preset compression ratio threshold of 50%, so continue to compress.
[0152] After the first audio data undergoes the third compression process, the compression ratio of the corresponding third compressed audio data is 70%, which reaches the preset compression ratio threshold of 50%, so stop the compression.
[0153] Determine the third compressed audio data with a compression ratio of 70% as the first audio data of the second encoding format to be sent to the server.
[0154] It should be noted that the above method of converting the first audio data from the first encoding format to the second encoding format through at least one compression process is only for example.
[0155] In some embodiments, other compression methods or other encoding format conversion methods can also be used to directly specify the conditions of the compressed audio data, perform only one compression or format conversion process, and directly obtain the first audio data of the second encoding format that meets the requirements. This embodiment does not limit this.
[0156] In some embodiments, in addition to determining whether the network environment meets the weak network requirements according to the reply feedback situation of the server within the specified period, other information can also be collected to determine the current network status, and then decide the encoding format used when the client sends the first audio data to the server.
[0157] Optionally, when obtaining the first audio data, obtain the signal strength information and location information of the terminal corresponding to the client.
[0158] The signal strength information is used to indicate the signal strength of the environment where the terminal is located, and the location information is used to indicate the geographical location coordinates of the terminal corresponding to the client, reflecting the type of environment where the terminal is located (for example, indoor, outdoor, underground).
[0159] Determine the network status information based on the location information and signal strength information, and the network status information is used to indicate the network status when the client transmits data.
[0160] Convert the location information and signal strength information into corresponding information values based on a preset rule, and assign corresponding network quality weights Q1 and Q2.
[0161] Obtain the network status weight based on the weighted operation result between the information value and the network quality weight, and use it as the network status information to quantify the communication quality of the network environment where the terminal is located, and then reflect the degree of matching between the current network environment and the weak network.
[0162] Exemplarily, obtain a preset information value conversion table, and the conversion table includes a first conversion relationship between location information and information values, and a second conversion relationship between signal strength information and information values.
[0163] Location information: It is a shopping mall on the second basement floor (altitude -10 meters, GPS signal strength -150dBm). Looking up the table, the corresponding environment type is a high-attenuation indoor environment, and the converted corresponding information value is 0.9.
[0164] Among them, the GPS (Global Positioning System) signal strength refers to the power level of the radio signal received by the receiver (such as a mobile phone or other terminal) from the GPS satellite, and is usually used to measure the positioning availability and accuracy.
[0165] The GPS signal strength of -150dBm means that the received GPS signal power is 10^-15 watts (0.000000000000001 watt), which belongs to an extremely weak signal.
[0166] The signal strength information includes: RSRP = -110dBm (the information value obtained by looking up the table is 0.7) and SINR = 4dB (the information value obtained by looking up the table is 0.5). Take the average value to convert the whole into the corresponding information value:
[0167] It is (0.5 + 0.7) / 2 = 1.2 / 2 = 0.6.
[0168] Among them, RSRP (Reference Signal Received Power) is the average power of the base station reference signal received by terminals such as mobile phones in the 4G / 5G network, which is used to measure the signal coverage intensity.
[0169] RSRP = -110 dBm (decibel milliwatt) belongs to an extremely weak signal.
[0170] SINR (Signal-to-Interference-plus-Noise Ratio) represents the ratio of the useful signal to (interference signal + noise), which is used to measure the signal quality (anti-interference ability). SINR = 4 dB (decibel) indicates being in a high-interference environment.
[0171] The network quality weight Q1 corresponding to the location information is 0.6, and the network quality weight Q2 corresponding to the signal strength information is 0.4. The calculated network status information is 0.6 * 0.9 + 0.4 * 0.6 = 0.54 + 0.24 = 0.78.
[0172] After the client sends the first audio data to the server in the first n preset transmission cycles in the first encoding format, when the network status information within the first n preset transmission cycles meets the weak network requirements, starting from the (n + 1)-th preset transmission cycle, the client sends the first audio data to the server in the second encoding format.
[0173] Exemplarily, meeting the weak network requirements means that the value in the network status information is within the specified value range.
[0174] For example, if the specified value range is (0.6, 0.9), then when the network status information is 0.78 and meets the weak network requirements, starting from the (n + 1)-th preset transmission cycle, the client sends the first audio data to the server in the second encoding format.
[0175] To sum up, in the audio recognition method provided by this application, during the audio recognition process, by periodically sending the audio data to be recognized, it is determined whether the current network environment is a weak network environment with a low transmission rate according to the reply feedback situation of the server within the specified number of cycles. The client decides the encoding format used when sending the audio data to the server according to the network environment situation. It can convert the transmission format from the first encoding format to the second encoding format with a smaller data volume in a weak network environment, reduce the transmission data volume of the audio data, reduce the client information request volume, and thus improve the transmission efficiency of the audio data. It avoids the situation of information transmission timeout due to a poor weak network environment and the client not being able to receive the server feedback result in time, and improves the success rate of audio recognition and the efficiency of obtaining the audio recognition result.
[0176] Figure 3It is a flowchart of an audio recognition method provided by another exemplary embodiment of the present application. This method is jointly executed by the client 310 and the server 320, and includes the following steps.
[0177] S311, the client obtains the first audio data.
[0178] Among them, the client can provide audio recognition function, and the first audio data is the data to be recognized by audio.
[0179] S312, the client sends the first audio data to the server in the first encoding format according to the preset sending period.
[0180] Among them, the first audio data is used to instruct the server to perform audio recognition.
[0181] Exemplarily, taking the scene of identifying a song by listening as an example, the audio recognition process executed by the server after receiving the first audio data is to determine the relevant information of the song corresponding to the first audio data, including but not limited to: song name, singer name, lyric information, etc.
[0182] The process that the client obtains the first audio data and sends the first audio data to the server is carried out in real time.
[0183] That is to say, the client continuously obtains the first audio data, and every time the preset sending period elapses, it sends the first audio data with different durations (the duration increases) to the server.
[0184] Among them, the first audio data in the first encoding format is audio data in Pulse Code Modulation (PCM) format. That is to say, sending the first audio data to the server in the first encoding format actually means sending PCM audio data to the server.
[0185] Among them, the PCM audio data is represented as a continuous byte array, in which the quantized amplitude values of each sampling point are stored.
[0186] That is to say, each value stored in the array refers to the amplitude value of each sampling point, representing the instantaneous amplitude of the sound wave at a certain moment.
[0187] A sampling point refers to a time node at which the continuous analog audio signal is sampled at a fixed time interval during the process of converting it into a discrete digital signal.
[0188] Optionally, the client sends the first audio data to the server in the first encoding format according to the preset sending period through the first communication interface.
[0189] The first communication interface is a data interaction interface between the client and the server, and is the transmission channel corresponding to the first audio data in the first encoding format.
[0190] S321. The server receives the first audio data in the first encoding format, and recognizes the received first audio data in the first encoding format to obtain a recognition result.
[0191] Optionally, the server receives the first audio data in the first encoding format through the first communication interface. For the first audio data in the same encoding format, the communication interface used by the client to send the first audio data is the same as the communication interface used by the server to receive the first audio data.
[0192] When the server receives the first audio data through the first communication interface, the server directly determines that the encoding format of the first audio data is the PCM format according to the type of the first communication interface.
[0193] The recognition result is used to describe the audio information of the first audio data.
[0194] For example, if the first audio data is the audio obtained by recording a certain song, the recognition result includes at least one of the following information of the song: the song name, the name of the singer who sings the song, the songwriters of the song, etc.
[0195] Optionally, the recognition process is as follows: The server receives the first audio data in the first encoding format, extracts the audio fingerprint of the first audio data in the first encoding format to obtain the first audio fingerprint feature.
[0196] Exemplarily, the first audio data in the first encoding format refers to PCM audio data. The server preprocesses the PCM audio data, converts the original PCM audio data (such as 16-bit signed integers) into a unified sampling rate (such as 8 kHz or 16 kHz) and mono (if it is stereo, it needs to be mixed), and performs audio amplitude normalization (for example, making its range belong to -1 to 1) to eliminate the volume difference.
[0197] Next, the server performs frame processing on the preprocessed PCM audio data, usually 20 - 40 milliseconds per frame, with 50% overlap, and then converts each frame of the signal into a spectrogram through the short-time Fourier transform method to reveal the change of frequency over time.
[0198] In some embodiments, in order to be closer to the human ear perception, the spectrum may be converted into the Mel scale or the logarithmic energy spectrum to highlight the significant frequency bands.
[0199] Local energy peaks are detected in the spectrum through a feature extraction algorithm, and these peak points (such as 500 Hz and 1 kHz at a certain moment) constitute the key features of the audio, and a feature vector is formed by pairing adjacent peaks (such as combining two frequency points and their time difference).
[0200] Convert the feature vector (e.g., encoding the frequency difference and time difference into integers) into a compact fingerprint hash value through a hash function, that is, convert it into a hash value with a fixed length, obtaining a hash sequence as the first audio fingerprint feature.
[0201] Optionally, obtain a fingerprint feature database, which contains at least two candidate audio fingerprint features respectively corresponding to at least two audio data.
[0202] Among them, the fingerprint feature database is a pre-prepared database. The process of obtaining the candidate audio fingerprint features in the fingerprint feature database refers to the above process of obtaining the first audio fingerprint feature, which will not be elaborated here.
[0203] Perform matching analysis based on the first audio fingerprint feature and the fingerprint feature database to obtain a matching result, which contains the matching degrees between the first audio fingerprint feature and at least two candidate audio fingerprint features respectively. Obtain the recognition result based on the matching result.
[0204] Exemplarily, the process of performing matching analysis based on the first audio fingerprint feature and the fingerprint feature database means that by comparing the hash sequence in the first audio fingerprint feature with the hash sequences of at least two candidate audio fingerprint features in the fingerprint feature database, calculate the similarity respectively, and determine the matching result based on the magnitude of the similarity value.
[0205] Exemplarily, for the i-th fingerprint feature among at least two candidate audio fingerprint features, calculate the similarity between the i-th fingerprint feature and the first audio fingerprint feature to obtain the i-th similarity, where i is a positive integer.
[0206] In the case where the i-th similarity meets the requirements of the preset matching pair, determine the i-th fingerprint feature as the target fingerprint feature.
[0207] Obtain an audio database, which contains at least two audio data and corresponding audio information.
[0208] Index the corresponding target audio data and corresponding target audio information from the audio database based on the target fingerprint feature.
[0209] Return the target audio information as the recognition result to the client.
[0210] For example, assume that there are fingerprint features A1 and B1 corresponding to audio A and audio B respectively in the fingerprint feature database (which are respectively represented by hash sequences H1 and H2 composed of several hash values).
[0211] After obtaining the first audio fingerprint feature, the server extracts the hash sequence H0 therefrom and compares the number of matching hashes S1 between the hash sequence H0 and the hash sequence H1, and the number of matching hashes S2 between the hash sequence H0 and the hash sequence H2.
[0212] Determine the fingerprint feature corresponding to the larger value among the hash quantities S1 and S2 as the candidate audio fingerprint feature that matches the first audio fingerprint feature, determine the audio data corresponding to the candidate audio fingerprint feature as the audio data that matches the first audio data, and obtain the corresponding audio information based on the audio data as the recognition result and return it to the client.
[0213] Exemplarily, in the specific matching process, the server will compare each item of the hash sequence H0 of the first audio data (assumed to be [0x3A1, 0x7B2, 0x5C3...]) with the corresponding positions of the hash sequences corresponding to each candidate fingerprint feature in the fingerprint feature database.
[0214] For example: If the 2-5th hash values of H0 (0x7B2, 0x5C3, 0x9D4, 0x2E5) are exactly the same as the 8-11th bits in the H1 sequence of audio A, while only the 3rd bit (0x5C3) of H0 matches the H2 sequence of audio B.
[0215] If the consecutive hashes that match maintain the same interval on the time axis (such as an interval of 200 ms between every two hash values), it is determined that there are more matching segments between H0 and H1 (4 consecutive matches vs 1 isolated match), and the timing rules are the same, and the hash quantity S1 is greater than the hash quantity S2. Finally, it is confirmed that the first audio data corresponds to audio A.
[0216] S322, the server returns the recognition result to the client.
[0217] Among them, after the server determines the candidate audio fingerprint feature corresponding to the first audio fingerprint feature, it will recall the target audio corresponding to the candidate audio fingerprint feature from the audio database of the server as the return result.
[0218] Among them, the return result includes the audio information of the target audio. For example, if the target audio is song A, the audio information includes the song name, singer name, album name to which the song belongs, lyrics, duration, etc. of song A.
[0219] S313, the client obtains the reply feedback situation of the server for the first audio data sent in the previous n preset sending cycles based on the recognition result.
[0220] Among them, in the case where the number of reply feedbacks for the first audio data sent by the server in the previous n preset sending cycles does not reach the first number threshold, it is determined that the reply feedback meets the weak network requirements.
[0221] Among them, meeting the weak network requirements means that the current network environment for information transmission between the client and the server is a weak network environment.
[0222] That is, when the server returns the recognition result to the client, the network environment is not sufficient to support the client to receive the recognition result sent by the server in a timely manner every time.
[0223] For example, if the first number threshold is n, and the number of returned results received within the first n preset cycles does not reach n times, it indicates that the current network environment is a weak network environment and meets the weak network requirements.
[0224] Exemplarily, the client maintains a weak network identifier, and the value of the weak network identifier is used to indicate whether the current network environment is a weak network environment.
[0225] The client obtains the weak network identifier based on the reception situation of the recognition result, and the value of the weak network identifier is used to describe the type of the current network environment.
[0226] In response to the number of replies not reaching the first number threshold, it is determined that the value of the weak network identifier is the first value false, and the current network environment is not a weak network environment.
[0227] In response to the number of replies reaching the first number threshold, it is determined that the value of the weak network identifier is the second value true, and the current network environment is a weak network environment.
[0228] Among them, the first value is different from the second value.
[0229] S314. When the reply feedback of the first audio data sent by the server for the first n preset sending cycles meets the weak network requirements, the client starts from the (n + 1)-th preset sending cycle and sends the first audio data to the server in the second encoding format.
[0230] Optionally, starting from the (n + 1)-th preset sending cycle, the first audio data is sent to the server in the second encoding format through the second communication interface.
[0231] Among them, the first communication interface and the second communication interface are different interfaces for data interaction between the client and the server.
[0232] Among them, when the first audio data is the same, the amount of encoded data result corresponding to the second encoding format is less than the amount of encoded data result corresponding to the first encoding format, and n is a positive integer.
[0233] Exemplarily, the first audio data in the second encoding format refers to audio data in the Moving Picture Experts Group Audio Layer 3 (MP3) format. That is, sending the first audio data to the server in the second encoding format actually means sending MP3 audio data to the server.
[0234] When the server receives the first audio data through the second communication interface, the server directly determines that the encoding format of the first audio data is the MP3 format according to the type of the second communication interface.
[0235] Optionally, when the weak network flag takes the value of true, the client performs format conversion on the first audio data, changing the format of the first audio data from the first encoding format to the second encoding format.
[0236] Exemplarily, the client performs compression processing on the first audio data to obtain compressed audio data that meets the requirements of a preset compression ratio, and sends the compressed audio data to the server through the second communication interface.
[0237] S323, the server receives the first audio data in the second encoding format.
[0238] Among them, the server receives the first audio data in the second encoding format through the second communication interface.
[0239] The server determines whether it is necessary to perform decompression processing on the first audio data before recognition according to the type of the second communication interface. That is, the server and the client directly determine that the encoding format of the first audio data is the MP3 format and decompression processing is required.
[0240] S324, the server performs decompression and recognition processing on the first audio data in the second encoding format to obtain a recognition result.
[0241] Optionally, the first audio data in the second encoding format is decompressed to obtain the first audio data in the first encoding format.
[0242] Audio fingerprint extraction is performed on the first audio data to obtain the first audio fingerprint feature. A fingerprint feature database is obtained, and the fingerprint feature database contains at least two candidate audio fingerprint features respectively corresponding to at least two audio data.
[0243] Matching analysis is performed based on the first audio fingerprint feature and the fingerprint feature database to obtain a matching result. The matching result includes the matching degrees between the first audio fingerprint feature and at least two candidate audio fingerprint features respectively. The recognition result is obtained based on the matching result.
[0244] Exemplarily, for the i-th fingerprint feature among at least two candidate audio fingerprint features, the similarity between the i-th fingerprint feature and the first audio fingerprint feature is calculated to obtain the i-th similarity, where i is a positive integer.
[0245] When the i-th similarity meets the requirements of a preset matching pair, the i-th fingerprint feature is determined as the target fingerprint feature.
[0246] An audio database is obtained, where the audio database contains at least two audio data and corresponding audio information.
[0247] Based on the target fingerprint feature, the corresponding target audio data and the corresponding target audio information are indexed from the audio database.
[0248] The target audio information is returned to the client as an updated recognition result.
[0249] The target audio data and the first audio data conform to a preset audio source relationship, and the audio source relationship includes at least one of the following relationships:
[0250] 1. The first audio data and the target audio data are from the same audio source: for example, the first audio data and the target audio data are partial segments of the same audio work;
[0251] 2. The target audio data includes first audio data: for example, the first audio data is a partial segment of the target audio data;
[0252] 3. There is an association between the first audio data and the target audio data: for example, the first audio data and the target audio data are different interpretation versions of the same audio work. Taking song A as an example, the first audio data is the live singing version of song A, and the target audio data is the studio version of song A.
[0253] The process of the server identifying the first audio data in step S324 and step S321 is the same, except that after the server receives the first audio data in the second encoding format, it needs to convert the second encoding format into the first encoding format before performing the audio fingerprint feature extraction and matching process, which will not be described here.
[0254] S325, the server returns the recognition result to the client.
[0255] The server returns the recognition result to the client through the second communication interface.
[0256] S315, the client receives the recognition result and displays the recognition result through a display interface.
[0257] Exemplarily, the recognition result indicates that the first audio data corresponds to song A, and information about song A is displayed on the display interface, such as the song title and the name of the singer who sang the song.
[0258] In some embodiments, if the first audio data comes from audio that is being played and is still being played, the client will align the current playback progress according to the recognition result returned by the server and play song A synchronously.
[0259] In summary, in the audio recognition method provided by the present application, during the audio recognition process, by periodically sending the audio data to be recognized, it is determined whether the current network environment is a weak network environment with a low transmission rate according to the reply feedback of the server within a specified number of cycles, and the client determines the encoding format used when sending the audio data to the server according to the network environment. It can convert the transmission format from the first encoding format to the second encoding format with a smaller data volume in the case of a weak network environment, reduce the transmission data volume of the audio data, reduce the information request volume of the client, and thus improve the transmission efficiency of the audio data. It avoids the situation of information transmission timeout caused by the poor weak network environment and the client not being able to receive the server feedback result in time, and improves the success rate of audio recognition and the efficiency of obtaining the audio recognition result.
[0260] Figure 4 FIG. is a flowchart of a song recognition method provided by an exemplary embodiment of the present application, which is executed by a client and includes the following steps.
[0261] S400, receive an audio recognition operation and start recording audio.
[0262] Schematically, referring to Figure 5 , Figure 5 is a schematic diagram of a song recognition interface of a client.
[0263] The song recognition interface 500 includes a target control 510 for triggering the audio recognition function, and option controls corresponding to different audio recognition types respectively:
[0264] (1) A humming recognition control 501 for collecting the first audio data from the sound signal of the user's vocal humming;
[0265] (2) A video song recognition control 502 for recording the audio of the video being played by the client to obtain the first audio data;
[0266] (3) A link song recognition control 503 for recording the audio of the multimedia content stored in the link to obtain the first audio data;
[0267] (4) A Live recognition control 504 for recording the background sound of the live environment to obtain the first audio data.
[0268] Among them, when the user triggers audio recognition, they can first trigger one of the above four option controls and then trigger the target control 510 to start recording audio to obtain the first audio data, or directly trigger the target control 510 to start recording audio to obtain the first audio data. This embodiment does not limit this.
[0269] Among them, for different audio recognition scenarios, the server uses different processing methods when recognizing the collected first audio data, which can improve the accuracy of the recognition result and the success rate of audio recognition.
[0270] Exemplarily, after the user triggers the humming recognition control 501, it corresponds to the humming recognition scenario.
[0271] In the humming recognition scenario, first, noise reduction processing is performed on the vocal humming recorded by the user to eliminate interference factors such as ambient noise and breathing sounds.
[0272] Subsequently, the fundamental frequency detection algorithm is used to extract the fundamental frequency sequence of the humming audio and convert it into normalized pitch-time series data.
[0273] Regarding the possible pitch deviation and rhythm deviation in the user's humming, the dynamic time warping algorithm is used to stretch and compress the extracted melody features on the time axis to achieve the optimal alignment and matching with the original melody features of the songs in the music library.
[0274] Exemplarily, after the user triggers the video music recognition control 502, it corresponds to the video music recognition scenario. In the video music recognition scenario, the sound source separation technology is adopted, and the background music in the video is separated from the sound tracks such as dialogue and sound effects through a deep neural network.
[0275] The system extracts the Mel-frequency cepstral coefficient features from the separated music signal and focuses on retaining the energy distribution features in the frequency band from 200 Hz to 5000 Hz.
[0276] For the 128 kbps bitrate compressed audio commonly used on video platforms, the system uses an anti-compression hashing algorithm to generate a 64-dimensional feature vector for matching.
[0277] Exemplarily, after the user triggers the link music recognition control 503, it corresponds to the link music recognition scenario. In the link music recognition scenario, the audio stream data is directly obtained by parsing the multimedia link, and the fingerprint algorithm is used to generate a 256-bit music fingerprint.
[0278] The standard feature database corresponding to the audio library (including the music fingerprints of multiple audios) is obtained, and the approximate nearest neighbor search algorithm is used to complete the retrieval and matching of the fingerprints.
[0279] Exemplarily, after the user triggers the Live recognition control 504, it corresponds to the Live recognition scenario. The four-microphone array is used to collect the audio signal, and the main lobe direction is pointed to the Live stage sound source through the beamforming algorithm. The short-time Fourier transform features of each frame of audio are calculated in real time, and a 128-dimensional spectral feature vector of every 2-second audio segment is extracted in a sliding window manner.
[0280] Regarding the non-linear distortion caused by the Live sound reinforcement system, a pre-trained deep neural network is used for feature compensation.
[0281] During the recognition process, the current recorded duration will be displayed in the song recognition interface 500. For example, Figure 5 The interface content when the recorded audio is 9 seconds is shown in, and a prompt message is displayed: "Get closer to the sound source for more accurate recognition."
[0282] That is, the closer to the sound source during audio recording, the higher the efficiency and accuracy of server recognition.
[0283] The combination trigger of the option control for different audio recognition types and the target control 510 can improve the accuracy and efficiency of the audio recognition result.
[0284] S401, Transmit the first 3 seconds of PCM audio to the server.
[0285] PCM audio is the first audio data. PCM audio data is represented as a continuous byte array, which stores the quantization amplitude value of each sampling point.
[0286] That is, each value stored in the array refers to the amplitude value of each sampling point, representing the instantaneous amplitude of the sound wave at a certain moment.
[0287] A sampling point refers to the time node at which the continuous analog audio signal is sampled at a fixed time interval during the conversion to a discrete digital signal.
[0288] In this embodiment, taking the recording process lasting 15 seconds as an example, that is, the maximum duration of the first audio data is 15 seconds.
[0289] Among them, with a preset transmission period of 3 seconds, the first audio data is collected in real time and sent to the server based on the preset transmission period. That is, a total of 15 / 3 = 5 times of the first audio data are sent.
[0290] Each time the PCM audio is sent, the duration of the sent PCM audio gradually increases.
[0291] S402, Determine whether the song recognition result is returned by the first communication interface.
[0292] The client uses the first communication interface to send PCM audio to the server, and the server also returns the song recognition result through the first communication interface after recognizing based on the PCM audio.
[0293] Among them, the song recognition result includes the result of recognizing the song to which the PCM audio belongs.
[0294] S403, Receive the song recognition result returned by the server.
[0295] If the server can return the song recognition result, the client receives the song recognition result.
[0296] S404, Transmit the PCM audio in the first 6 seconds to the server.
[0297] When the recorded audio duration reaches 6 seconds, the second preset transmission cycle is reached, and the PCM audio from the 1st second to the 6th second is sent to the server through the first communication interface.
[0298] S405, Determine whether the first communication interface returns a music recognition result.
[0299] Refer to step S402.
[0300] S406, Receive the music recognition result returned by the server.
[0301] If the server can return a music recognition result, the client receives the music recognition result.
[0302] S407, Determine whether results are returned for both of the previous two audio transmission requests.
[0303] Determining whether results are returned for both is to determine whether the network environment between the server and the client can support the client to perform audio recognition by transmitting audio data in PCM format.
[0304] S408, Meeting the judgment requirements, transmit the PCM audio in the first 9 seconds to the server, and determine that the weak network flag is false.
[0305] If the information is transmitted in a timely manner and the client can receive the recognition result returned by the server each time, it means that the current network environment is not a weak network environment, the weak network flag is false, and PCM format data can still be used for transmission.
[0306] S409, Not meeting the judgment requirements, compress the PCM audio in the first 9 seconds into MP3 audio and transmit it to the server, and determine that the weak network flag is true.
[0307] If the information transmission times out and the recognition result returned by the server cannot be received each time, it means that the current network environment is a weak network environment, the weak network flag is true, which means that when using PCM format data for transmission, the audio recognition process will fail, and the PCM data needs to be compressed to obtain MP3 format audio data with a smaller data volume, and the MP3 audio is sent to the server.
[0308] Among them, the transmitted MP3 audio refers to the audio from the 1st second to the 9th second.
[0309] S410, Determine whether the communication interface returns a music recognition result.
[0310] If the client sends audio data using the first communication interface, determine whether the first communication interface returns a song recognition result; if the client sends audio data using the second communication interface, determine whether the second communication interface returns a song recognition result.
[0311] S411, Receive the song recognition result returned by the server.
[0312] If the server can return a song recognition result, the client receives the song recognition result.
[0313] S412, Determine whether the weak network flag is true, and transmit the first 12 seconds of audio in the corresponding format to the server.
[0314] If the weak network flag is true, it means to transmit the first 12 seconds of audio in MP3 format to the server; if the weak network flag is false (not true), it means to transmit the first 12 seconds of audio in PCM format to the server.
[0315] S413, In the case where the weak network flag is true, compress the first 12 seconds of PCM audio into MP3 audio and transmit it to the server.
[0316] Among them, the first 12 seconds of audio transmitted refers to the audio from the 1st second to the 12th second counted from the start time of recording.
[0317] S414, Determine whether the second communication interface returns a song recognition result.
[0318] The client uses the second communication interface to send MP3 audio to the server, and the server also returns the song recognition result through the second communication interface after recognizing based on the MP3 audio.
[0319] Among them, the song recognition result includes the result of recognizing the song to which the MP3 audio belongs.
[0320] S415, Receive the song recognition result returned by the server.
[0321] If the server can return a song recognition result, the client receives the song recognition result.
[0322] S416, Determine whether the weak network flag is true, and transmit the first 15 seconds of audio in the corresponding format to the server.
[0323] If the weak network flag is true, it means to transmit the first 15 seconds of audio in MP3 format to the server; if the weak network flag is false (not true), it means to transmit the first 15 seconds of audio in PCM format to the server.
[0324] S417, In the case where the weak network flag is true, compress the first 15 seconds of PCM audio into MP3 audio and transmit it to the server.
[0325] Among them, the first 15 seconds of the transmitted audio refers to the audio from the 1st second to the 15th second counted from the start time of recording.
[0326] S418, determine whether the second communication interface returns a music recognition result.
[0327] Refer to step S414.
[0328] S419, receive the music recognition result returned by the server.
[0329] If the server can return a music recognition result, the client receives the music recognition result.
[0330] Among them, the music recognition results returned by the server in steps S403, S406, S411, S415, and step S419 are not necessarily exactly the same. The longer the audio duration and the clearer the audio quality used by the server during recognition, the more accurate the music recognition result.
[0331] Exemplarily, the music recognition result obtained by the server based on 15 seconds of audio in step S419 is determined as the final music recognition result.
[0332] In summary, the audio recognition method provided by this application, during the process of audio recognition, periodically sends the audio data to be recognized, determines whether the current network environment is a weak network environment with a low transmission rate according to the reply feedback of the server within a specified number of cycles, and the client determines the encoding format used when sending audio data to the server according to the network environment. It can convert the transmission format from the first encoding format to the second encoding format with a smaller data volume in a weak network environment, reduce the transmission data volume of the audio data, reduce the amount of information requests from the client, and thus improve the transmission efficiency of the audio data. Avoid the situation of information transmission timeout caused by a poor weak network environment and the client not being able to receive the server feedback result in a timely manner, and improve the success rate of audio recognition and the efficiency of obtaining the audio recognition result.
[0333] Figure 6 It is a structural block diagram of an audio recognition device provided by an exemplary embodiment of this application. As Figure 6 shown, this device includes the following parts.
[0334] An acquisition module 610, configured to acquire first audio data, where the first audio data is the data to be subjected to audio recognition;
[0335] A sending module 620, configured to send the first audio data to the server in the first encoding format according to a preset sending period, where the first audio data is used to instruct the server to perform audio recognition;
[0336] The sending module 620 is further configured to, when the reply feedback of the first audio data sent by the server for the first n preset sending cycles meets the weak network requirements, start from the (n + 1)-th preset sending cycle, and send the first audio data to the server in a second encoding format, where, for the same first audio data, the data volume of the encoding data result corresponding to the second encoding format is smaller than that of the encoding data result corresponding to the first encoding format, and n is a positive integer.
[0337] In an alternative embodiment, the sending module 620 is further configured to determine that the reply feedback meets the weak network requirements when the number of reply times of the reply feedback of the first audio data sent by the server for the first n preset sending cycles does not reach the first threshold; and start from the (n + 1)-th preset sending cycle, and send the first audio data to the server in the second encoding format.
[0338] In an alternative embodiment, the sending module 620 is further configured to send the first audio data to the server in the first encoding format according to the preset sending cycle through a first communication interface;
[0339] The sending module 620 is further configured to start from the (n + 1)-th preset sending cycle, and send the first audio data to the server in the second encoding format through a second communication interface; where the first communication interface and the second communication interface are different interfaces for data interaction between the client and the server.
[0340] In an alternative embodiment, the sending module 620 is further configured to start from the (n + 1)-th preset sending cycle, perform format conversion on the first audio data, and change the format of the first audio data from the first encoding format to the second encoding format; and send the first audio data to the server in the second encoding format.
[0341] In an alternative embodiment, the sending module 620 is further configured to perform at least one compression process on the first audio data to obtain at least one compressed audio data; obtain the compression ratio of each of the at least one compressed audio data with respect to the first audio data; stop the compression process in response to the compression ratio corresponding to the current compressed audio data obtained by the current compression process meeting the preset format conversion requirements; and determine the current compressed audio data as the first audio data corresponding to the second encoding format.
[0342] In an alternative embodiment, the sending module 620 is further configured to perform an i-th compression process on the first audio data to obtain an i-th compressed audio data; calculate a data volume difference between the i-th compressed audio data and the first audio data to obtain an i-th differential audio data, where i is a positive integer; and obtain an i-th compression rate based on a ratio between the i-th differential audio data and the first audio data.
[0343] In an alternative embodiment, the sending module 620 is further configured to stop the compression process when the (k - 1)-th compression rate obtained after the (k - 1)-th compression process does not reach a preset compression rate threshold and the k-th compression rate obtained after the k-th compression process reaches the preset compression rate threshold; where k is a positive integer; and determine the compressed audio data corresponding to the k-th compression rate as the first audio data corresponding to the second coding format.
[0344] In an alternative embodiment, the obtaining module 610 is further configured to collect an audio data segment based on the preset sending period to obtain the first audio data, where a duration of the audio data segment matches a period duration of the preset sending period.
[0345] The sending module 620 is further configured to, in the n-th preset sending period, integrate the audio data segments respectively collected in the previous n preset sending periods to obtain an n-th integrated audio data, and use the n-th integrated audio data as the first audio data to be sent in the n-th preset sending period; and send the n-th integrated audio data to the server in the first coding format in the n-th preset sending period.
[0346] In an alternative embodiment, the obtaining module 610 is further configured to, when obtaining the first audio data, obtain signal strength information and location information of the terminal corresponding to the client, where the signal strength information is used to indicate a signal strength of the environment where the terminal is located; and determine network state information based on the location information and the signal strength information, where the network state information is used to indicate a network state when the client transmits data.
[0347] The sending module 620 is further configured to, when the network state information in the previous n preset sending periods meets the weak network requirements, start from the (n + 1)-th preset sending period, and send the first audio data to the server in the second coding format.
[0348] In summary, during the audio recognition process, the audio recognition device provided in this application periodically sends the audio data to be recognized, determines whether the current network environment is a weak network environment with a low transmission rate based on the response feedback of the server within a specified number of cycles, and the client determines the encoding format used when sending audio data to the server according to the network environment. It can convert the transmission format from the first encoding format to the second encoding format with a smaller data volume in the case of a weak network environment, reduce the transmission data volume of the audio data, reduce the information request volume of the client, and thus improve the transmission efficiency of the audio data. It can avoid the situation of information transmission timeout caused by the poor weak network environment and the client not being able to receive the server feedback result in time, and improve the success rate of audio recognition and the efficiency of obtaining the audio recognition result.
[0349] It should be noted that: for the audio recognition device provided in the above embodiment, only the above division of each functional module is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the audio recognition device provided in the above embodiment and the audio recognition method embodiment belong to the same concept, and the specific implementation process can be seen in the method embodiment, which will not be elaborated here.
[0350] Figure 7 The block diagram of the computer device 700 provided by an exemplary embodiment of this application is shown. The computer device 700 can be: a smart phone, a tablet computer, a Moving Picture Experts Group Audio Layer III player (MP3), a Moving Picture Experts Group Audio Layer IV (MP4) player, a notebook computer or a desktop computer.
[0351] The computer device 700 may also be called by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc.
[0352] Generally, the computer device 700 includes: a processor 701 and a memory 702.
[0353] The processor 701 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc.
[0354] The processor 701 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA).
[0355] The processor 701 may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the central processing unit (CPU); the coprocessor is a low-power processor used to process data in the standby state.
[0356] In some embodiments, the processor 701 may be integrated with a graphics processing unit (GPU), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen.
[0357] In some embodiments, the processor 701 may further include an artificial intelligence (AI) processor, and the AI processor is used to process computational operations related to machine learning.
[0358] The memory 702 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory.
[0359] The memory 702 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices.
[0360] In some embodiments, the non-transitory computer-readable storage media in the memory 702 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 701 to implement the audio recognition method provided in the method embodiments of the present application.
[0361] In some embodiments, the computer device 700 may further include some other components 703, and the type and quantity of the other components 703 may be selected based on the functional requirements of the computer device 700.
[0362] Those skilled in the art can understand that Figure 7 the structure shown in does not limit the computer device 700, and it may include more or fewer components than shown in the figure, or combine certain components, or adopt different component arrangements.
[0363] Optionally, the computer-readable storage medium may include: Read Only Memory (ROM), Random Access Memory (RAM), Solid State Drives (SSD), optical discs, etc.
[0364] Among them, the random access memory may include Resistance Random Access Memory (ReRAM) and Dynamic Random Access Memory (DRAM). The serial numbers of the embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.
[0365] The embodiments of the present application further provide a computer device, which includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the audio recognition method as described in any one of the embodiments of the present application above.
[0366] The embodiments of the present application further provide a computer-readable storage medium, in which at least one instruction, at least one program, a code set, or an instruction set is stored, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the audio recognition method as described in any one of the embodiments of the present application above.
[0367] The embodiments of the present application further provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium.
[0368] The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the audio recognition method as described in any one of the above embodiments.
[0369] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and the storage medium mentioned above can be a read-only memory, a magnetic disk, an optical disc, etc.
[0370] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. An audio recognition method, characterized in that, The method includes: Obtaining first audio data, where the first audio data is the data to be subjected to audio recognition; Sending the first audio data to a server in a first encoding format according to a preset sending period, where the first audio data is used to instruct the server to perform audio recognition; When the reply feedback from the server for the first audio data sent in the previous n preset sending periods meets the weak network requirements, starting from the (n + 1)-th preset sending period, sending the first audio data to the server in a second encoding format, where, when the first audio data is the same, the data volume of the encoding data result corresponding to the second encoding format is smaller than the data volume of the encoding data result corresponding to the first encoding format, and n is a positive integer.
2. The method according to claim 1, wherein The step of, when the reply feedback from the server for the first audio data sent in the previous n preset sending periods meets the weak network requirements, starting from the (n + 1)-th preset sending period, sending the first audio data to the server in a second encoding format includes: When the number of reply times of the reply feedback from the server for the first audio data sent in the previous n preset sending periods does not reach a first number threshold, determining that the reply feedback meets the weak network requirements; Starting from the (n + 1)-th preset sending period, sending the first audio data to the server in the second encoding format.
3. The method according to claim 1, wherein The step of sending the first audio data to the server in a first encoding format according to a preset sending period includes: Sending the first audio data to the server in the first encoding format according to the preset sending period through a first communication interface; The step of, starting from the (n + 1)-th preset sending period, sending the first audio data to the server in a second encoding format includes: Starting from the (n + 1)-th preset sending period, sending the first audio data to the server in the second encoding format through a second communication interface; Wherein, the first communication interface and the second communication interface are different interfaces for data interaction between the client and the server.
4. The method according to any one of claims 1 to 3, characterized in that, The step of, starting from the (n + 1)-th preset sending period, sending the first audio data to the server in a second encoding format includes: Starting from the (n + 1)-th preset sending period, performing format conversion on the first audio data, and changing the format of the first audio data from the first encoding format to the second encoding format; Sending the first audio data to the server in the second encoding format.
5. The method according to claim 4, wherein The step of performing format conversion on the first audio data and changing the format of the first audio data from the first encoding format to the second encoding format includes: Performing at least one compression process on the first audio data to obtain at least one compressed audio data; Obtaining the compression ratio of each of the at least one compressed audio data with respect to the first audio data; When the compression ratio corresponding to the current compressed audio data obtained by the current compression process meets a preset format conversion requirement, stopping the compression process; Determining the current compressed audio data as the first audio data corresponding to the second encoding format.
6. The method according to claim 5, wherein Obtaining the compression ratio of the at least one compressed audio data for the first audio data respectively includes: Performing an i-th compression process on the first audio data to obtain an i-th compressed audio data; Calculating the data volume difference between the i-th compressed audio data and the first audio data to obtain an i-th differential audio data, where i is a positive integer; Based on the ratio between the i-th differential audio data and the first audio data, obtaining an i-th compression ratio.
7. The method according to claim 5, wherein Responding to the compression ratio corresponding to the current compressed audio data obtained by the current compression process meeting the preset format conversion requirements and stopping the compression process includes: Stopping the compression process when the (k - 1)-th compression ratio obtained after the (k - 1)-th compression process does not reach the preset compression ratio threshold and the k-th compression ratio obtained after the k-th compression process reaches the preset compression ratio threshold; k is a positive integer; Determining the compressed audio data corresponding to the k-th compression ratio as the first audio data corresponding to the second coding format.
8. The method according to any one of claims 1 to 3, characterized in that Obtaining the first audio data includes: Collecting audio data segments based on the preset sending period to obtain the first audio data, where the duration of the audio data segments matches the period duration of the preset sending period; Sending the first audio data to the server in the first coding format according to the preset sending period includes: In the n-th preset sending period, integrating the audio data segments collected in the previous n preset sending periods to obtain an n-th integrated audio data as the first audio data sent in the n-th preset sending period; In the n-th preset sending period, sending the n-th integrated audio data to the server in the first coding format.
9. The method according to any one of claims 1 to 3, characterized in that The method further includes: When obtaining the first audio data, obtaining the signal strength information and location information of the terminal corresponding to the client, where the signal strength information is used to indicate the signal strength of the environment where the terminal is located; Determining network status information based on the location information and the signal strength information, where the network status information is used to indicate the network status when the client transmits data; After sending the first audio data to the server in the first coding format according to the preset sending period, it further includes: When the network status information in the previous n preset sending periods meets the weak network requirements, starting from the (n + 1)-th preset sending period, sending the first audio data to the server in the second coding format.
10. An audio recognition device, characterized in that, The apparatus includes: An obtaining module, configured to obtain first audio data, where the first audio data is data to be subjected to audio recognition; A sending module, configured to send the first audio data to the server in the first coding format according to the preset sending period, where the first audio data is used to instruct the server to perform audio recognition; The sending module is further configured to, when the reply feedback of the first audio data sent by the server for the first n preset sending periods meets the weak network requirements, start from the (n + 1)-th preset sending period, and send the first audio data to the server in a second coding format, where, when the first audio data is the same, the amount of data of the coding data result corresponding to the second coding format is less than the amount of data of the coding data result corresponding to the first coding format, and n is a positive integer.
11. A computer device, characterized in that, The computer device includes a processor and a memory, and at least one program is stored in the memory, and the at least one program is loaded and executed by the processor to implement the audio recognition method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, At least one program is stored in the storage medium, and the at least one program is loaded and executed by a processor to implement the audio recognition method according to any one of claims 1 to 9.
13. A computer program product, characterized in that, It includes a computer program, and when the computer program is executed by a processor, it implements the audio recognition method according to any one of claims 1 to 9.