Audio recognition method, apparatus and device

CN116312549BActive Publication Date: 2026-09-15BEIJING PHOENIX AUTO INTELLIGENCE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310215708.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-28
Publication Date
2026-09-15
Estimated Expiration
2043-02-28

AI Technical Summary

Benefits of technology

[0044] By splicing disjoint initial audio segments, a dramatic shift in intonation is created in the spliced ​​audio, resulting in target intonation features that are stronger than the initial intonation features of the unspliced ​​audio segments. These stronger target intonation features assist in audio recognition, thereby improving the accuracy of audio recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116312549B_ABST
    Figure CN116312549B_ABST
Patent Text Reader

Abstract

The application discloses an audio recognition method, device and equipment, and belongs to the technical field of computers. The method comprises the following steps: acquiring a plurality of initial audios, the plurality of initial audios corresponding to a same audio providing object, and object information of the audio providing object being unknown; splicing non-connected initial audios in the plurality of initial audios to obtain a plurality of spliced audios, the target intonation feature carried by the spliced audios being stronger than the initial intonation feature carried by the initial audios; acquiring a voiceprint of a reference audio, the object information corresponding to the reference audio being known; and determining an audio recognition result of the plurality of initial audios according to the voiceprint of the reference audio and the plurality of spliced audios. The non-connected initial audios are spliced, so that the target intonation feature carried by the spliced audios is stronger than the initial intonation feature. The target intonation feature with stronger characteristics is used to assist audio recognition, and the accuracy of audio recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to an audio recognition method, apparatus, and device. Background Technology

[0002] With the development of computer technology, more and more application scenarios are beginning to value information security and are using identity authentication technology to ensure it. For example, audio recognition is used to determine the identity of the audio provider for authentication purposes. Summary of the Invention

[0003] This application provides an audio recognition method, apparatus, and device, which can be used to determine the identity of an audio provider through audio recognition. The technical solution is as follows:

[0004] On one hand, embodiments of this application provide an audio recognition method, the method comprising:

[0005] Multiple initial audio files are obtained, each corresponding to the same audio provider object, and the object information of the audio provider object is unknown.

[0006] By splicing together the disjoint initial audios from the multiple initial audios, multiple spliced ​​audios are obtained, wherein the target intonation features carried by the spliced ​​audios are stronger than the initial intonation features carried by the initial audios.

[0007] The voiceprint of the reference audio is obtained, and the object information corresponding to the reference audio is known;

[0008] The audio recognition results of the multiple initial audios are determined based on the voiceprints of the multiple spliced ​​audios and the reference audio.

[0009] In one possible implementation, determining the audio recognition result of the plurality of initial audios based on the voiceprints of the plurality of concatenated audios and the reference audio includes:

[0010] Multiple first voiceprints are obtained from the multiple spliced ​​audios, and each spliced ​​audio corresponds to at least one first voiceprint. At least one first voiceprint in the first voiceprint corresponding to each spliced ​​audio is determined according to the target intonation features of the audio provider.

[0011] The plurality of first voiceprints are matched with the plurality of second voiceprints to obtain the audio matching results of the plurality of initial audios and the reference audio, wherein the second voiceprint is the voiceprint of the reference audio, and at least one of the plurality of second voiceprints is determined according to the target intonation feature corresponding to the reference audio.

[0012] The audio recognition results of the multiple initial audios are determined based on the audio matching results.

[0013] In one possible implementation, obtaining the multiple first voiceprints of the multiple spliced ​​audios includes:

[0014] For any concatenated audio, audio segmentation is performed on the concatenated audio to obtain at least one first audio. The at least one first audio contains an audio splicing point that includes the concatenated audio. The audio splicing point carries the target intonation feature of the audio providing object.

[0015] The voiceprints of each of the at least one first audio audios are extracted to obtain the first voiceprint.

[0016] In one possible implementation, the step of segmenting the concatenated audio to obtain at least one first audio segment includes:

[0017] Determine the length of the moving window used for audio segmentation and the step size of the moving window;

[0018] Based on the length and step size of the moving window, the audio of any spliced ​​audio is segmented with reference to the audio splicing point of any spliced ​​audio to obtain at least one first audio.

[0019] In one possible implementation, there are multiple moving windows for audio segmentation, and the length of any one of the multiple moving windows is different from the length of the other moving windows, wherein the other moving windows are the moving windows other than any one of the multiple moving windows.

[0020] In one possible implementation, matching the plurality of first voiceprints with a plurality of second voiceprints to obtain audio matching results between the plurality of initial audios and the reference audio includes:

[0021] Calculate the similarity between each of the plurality of first voiceprints and each of the plurality of second voiceprints;

[0022] The audio matching result is determined based on the similarity and matching threshold between each of the first voiceprints and each of the second voiceprints.

[0023] In one possible implementation, determining the audio matching result based on the similarity and matching threshold between the respective first voiceprints and the respective second voiceprints includes:

[0024] Based on the fact that the similarity between any first voiceprint and any second voiceprint is greater than the matching threshold, the audio matching result is determined to be a successful match.

[0025] Alternatively, based on the fact that the similarity between each of the first voiceprints and each of the second voiceprints is not greater than the matching threshold, the audio matching result is determined to be a matching failure.

[0026] In one possible implementation, determining the audio recognition result of the plurality of initial audios based on the audio matching result includes:

[0027] Based on the successful audio matching result, the object information corresponding to the reference audio is determined as the object information of the audio providing object, thus obtaining the audio recognition result of the multiple initial audios.

[0028] On the other hand, an audio recognition device is provided, the device comprising:

[0029] The acquisition module is used to acquire multiple initial audio files, which correspond to the same audio provider object, and the object information of the audio provider object is unknown.

[0030] The splicing module is used to splice unconnected initial audios from the plurality of initial audios to obtain a plurality of spliced ​​audios, wherein the target intonation features carried by the spliced ​​audios are stronger than the initial intonation features carried by the initial audios.

[0031] The acquisition module is also used to acquire the voiceprint of the reference audio, wherein the object information corresponding to the reference audio is known;

[0032] The determining module is used to determine the audio recognition result of the multiple initial audios based on the voiceprints of the multiple spliced ​​audios and the reference audio.

[0033] In one possible implementation, the determining module is configured to acquire multiple first voiceprints of the multiple concatenated audio files, where each concatenated audio file corresponds to at least one first voiceprint, and at least one of the first voiceprints corresponding to the concatenated audio file is determined based on the target intonation features of the audio provider; match the multiple first voiceprints with multiple second voiceprints to obtain an audio matching result between the multiple initial audio files and the reference audio file, where the second voiceprint is the voiceprint of the reference audio file, and at least one of the multiple second voiceprints is determined based on the target intonation features corresponding to the reference audio file; and determine the audio recognition result of the multiple initial audio files based on the audio matching result.

[0034] In one possible implementation, the determining module is configured to, for any concatenated audio, perform audio segmentation on the concatenated audio to obtain at least one first audio, wherein the at least one first audio contains an audio splicing point of the concatenated audio, and the audio splicing point carries the target intonation features of the audio providing object; and extract the voiceprint of each of the at least one first audio to obtain the first voiceprint.

[0035] In one possible implementation, the determining module is configured to determine the length of the moving window used for audio segmentation and the step size of the moving window; and to perform audio segmentation on any spliced ​​audio by referring to the audio splicing point of any spliced ​​audio, based on the length of the moving window and the step size of the moving window, to obtain the at least one first audio.

[0036] In one possible implementation, there are multiple moving windows for audio segmentation, and the length of any one of the multiple moving windows is different from the length of the other moving windows, wherein the other moving windows are the moving windows other than any one of the multiple moving windows.

[0037] In one possible implementation, the determining module is configured to calculate the similarity between each of the plurality of first voiceprints and each of the plurality of second voiceprints; and determine the audio matching result based on the similarity between each of the first voiceprints and each of the second voiceprints and a matching threshold.

[0038] In one possible implementation, the determining module is configured to determine the audio matching result as a successful match based on the fact that the similarity between any first voiceprint and any second voiceprint is greater than the matching threshold; or, to determine the audio matching result as a failed match based on the fact that the similarity between each first voiceprint and each second voiceprint is not greater than the matching threshold.

[0039] In one possible implementation, the determining module is used to determine the object information corresponding to the reference audio as the object information of the audio providing object based on the audio matching result being a successful match, thereby obtaining the audio recognition result of the multiple initial audios.

[0040] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor to enable the computer device to implement any of the audio recognition methods described above.

[0041] On the other hand, a computer-readable storage medium is also provided, wherein at least one computer program is stored in the computer-readable storage medium, the at least one computer program being loaded and executed by a processor to enable a computer to implement any of the audio recognition methods described above.

[0042] On the other hand, a computer program product or computer program is also provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to perform any of the audio recognition methods described above.

[0043] The technical solution provided in this application brings at least the following beneficial effects:

[0044] By splicing disjoint initial audio segments, a dramatic shift in intonation is created in the spliced ​​audio, resulting in target intonation features that are stronger than the initial intonation features of the unspliced ​​audio segments. These stronger target intonation features assist in audio recognition, thereby improving the accuracy of audio recognition. Attached Figure Description

[0045] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0046] Figure 1 This is a schematic diagram of an implementation environment provided in an embodiment of this application;

[0047] Figure 2 This is a flowchart of an audio recognition method provided in an embodiment of this application;

[0048] Figure 3 This is a flowchart of obtaining spliced ​​audio provided in an embodiment of this application;

[0049] Figure 4 This is a flowchart of obtaining a second voiceprint provided in an embodiment of this application;

[0050] Figure 5 This is a schematic diagram illustrating a process of segmenting and splicing audio according to an embodiment of this application;

[0051] Figure 6 This is a flowchart illustrating the segmentation and splicing of audio data provided in an embodiment of this application;

[0052] Figure 7This is a flowchart of obtaining a first voiceprint provided in an embodiment of this application;

[0053] Figure 8 This is a flowchart of a voiceprint matching method provided in an embodiment of this application;

[0054] Figure 9 This is a schematic diagram of the structure of an audio recognition device provided in an embodiment of this application;

[0055] Figure 10 This is a schematic diagram of the structure of a server provided in an embodiment of this application;

[0056] Figure 11 This is a schematic diagram of the structure of an audio recognition device provided in an embodiment of this application. Detailed Implementation

[0057] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0058] This application provides an audio recognition method. Please refer to the following embodiments. Figure 1 The diagram illustrates the implementation environment of the method provided in this application embodiment. This implementation environment may include: terminal 11 and server 12.

[0059] Optionally, this method can be executed independently by terminal 11 or server 12, or interactively by terminal 11 and server 12. For example, in the interactive execution process between terminal 11 and server 12, terminal 11 has an application installed capable of acquiring initial audio. After acquiring the initial audio, the application sends it to server 12. Server 12 then splices the initial audio using the audio recognition method provided in this embodiment to obtain spliced ​​audio. Server 12 determines the audio recognition result of the initial audio based on the spliced ​​audio. Optionally, server 12 sends the spliced ​​audio to terminal 11, and terminal 11 determines the audio recognition result of the initial audio based on the spliced ​​audio.

[0060] Optionally, terminal 11 can be any electronic product capable of human-computer interaction with the user through one or more methods such as a keyboard, touchpad, touchscreen, remote control, voice interaction, or handwriting device, such as PC (Personal Computer), mobile phone, smartphone, PDA (Personal Digital Assistant), wearable device, PPC (Pocket PC), tablet computer, smart car system, smart TV, smart speaker, etc. Server 12 can be a single server, a server cluster consisting of multiple servers, or a cloud computing service center. Terminal 11 and server 12 establish a communication connection through wired or wireless network.

[0061] Those skilled in the art should understand that the above-described terminal 11 and server 12 are merely examples. Other existing or future terminals or servers that are applicable to this application should also be included within the scope of protection of this application, and are hereby incorporated by reference.

[0062] This application provides an audio recognition method, which can be based on the above. Figure 1 In the implementation environment shown, this method can be executed by terminal 11, and the flowchart of the method is as follows. Figure 2 As shown, it includes steps 201-204.

[0063] In step 201, multiple initial audios are obtained, each corresponding to the same audio provider object, whose object information is unknown.

[0064] For example, the terminal acquires multiple audio data points, using one audio data point as an initial audio source, thus obtaining multiple initial audio sources. Alternatively, the terminal can acquire a segment of audio data, divide it into multiple segments based on whether the audio data meets segmentation criteria, and use one segment as an initial audio source. Meeting the segmentation criteria means that the audio data length is not less than a reference length. The reference length can be a value set based on experience, such as 10 minutes. If the audio data length is not less than the reference length, even if the terminal segments the audio data, the resulting initial audio sources will not be too short to extract the first voiceprint.

[0065] Whether the terminal acquires multiple audio data separately, uses each audio data as an initial audio, or segments a piece of audio data to obtain multiple initial audio, it needs to acquire audio data first. The terminal can acquire audio data in ways including but not limited to the following three methods.

[0066] Method 1: The terminal provides an audio input interface, and the audio provider inputs audio data based on the audio input interface.

[0067] The audio input interface can be any device capable of supporting audio capture, including but not limited to a microphone (MIC). Optionally, the microphone can be a built-in microphone or an external microphone. When an external microphone is connected, it can be achieved via a wired connection or a wireless connection. Wireless connections include, for example, Bluetooth or a wireless network connection.

[0068] This application does not limit the text read aloud by the audio provider when inputting audio data; it can be text read freely by the audio provider or text provided by the terminal. For example, the terminal randomly generates a piece of text and displays the generated text on the screen to assist the audio provider in inputting audio data. When the terminal inputs audio data through the audio input interface, it can also set input conditions. The terminal determines the end of audio data input based on the audio data input by the audio acquisition object meeting the input conditions.

[0069] For scenarios where multiple audio data points are needed as initial audio samples, the input condition can be a reference number of audio data points to be input. The audio provider needs to input the reference number of audio data points before the terminal determines that the audio data acquisition is complete. For scenarios where multiple initial audio samples are obtained by segmenting the same audio segment, the input condition can be a reference length of the audio data to be input. The audio provider needs to continuously input audio data until the length of the input audio data is not less than the reference length before the terminal determines that the audio data acquisition is complete.

[0070] The terminal determines the end of audio data input by, for example, automatically ending audio data acquisition and disconnecting from the audio acquisition device. Alternatively, the terminal provides an end control; based on whether the audio data input by the audio acquisition object meets the input conditions, the end control changes from a non-triggerable state to a triggerable state, allowing the audio acquisition object to terminate audio input by triggering the end control.

[0071] Method 2: Access the audio database and retrieve audio data from it.

[0072] Optionally, the audio database accessed by the terminal can be an audio database stored in the terminal's storage space, or it can be an audio database that the terminal has access to. The audio database internally stores audio data separately according to different audio providers. The terminal can access the audio database to retrieve audio data belonging to the same audio provider, or retrieve audio data that meets the segmentation criteria.

[0073] Method 3: Obtain the video and extract the audio data from it.

[0074] For example, the method by which the terminal acquires video is similar to the method for acquiring audio data in methods one and two described above, and can be referred to the relevant descriptions, which will not be repeated here. After acquiring the video, the terminal can use Python (a computer programming language) to extract the audio data from the video.

[0075] The terminal can choose any of the above-mentioned acquisition methods to acquire audio data, or it can use other acquisition methods to acquire audio data, and obtain the initial audio based on the audio data. Since there may be some interference information in the audio data, such as noise and silence, the terminal will process the audio data and use the processed audio data as the initial audio. For example, silence cancellation technology can be used to acquire speech audio. Silence cancellation technology, such as VAD (Voice Activity Detection), locates the start and end positions of the speech audio of the audio provider in the audio data using VAD, and performs audio segmentation based on the start and end positions to extract the speech audio from the audio data, thus obtaining the initial audio. Of course, the terminal can also use other silence cancellation technologies to extract speech audio; this application embodiment does not limit this approach.

[0076] Optionally, the lengths of the multiple initial audio files acquired by the terminal can be exactly the same, partially the same, or completely different; this embodiment does not limit this. Regardless of the method by which the terminal acquires the multiple initial audio files, the object information of the audio provider corresponding to the multiple initial audio files is unknown, and the object information of the audio provider needs to be determined through audio recognition. The object information may include, for example, the identity information, age information, and gender information of the audio provider.

[0077] In step 202, disjoint initial audio segments from multiple initial audio segments are concatenated to obtain multiple concatenated audio segments. The target intonation features carried by the concatenated audio segments are stronger than the initial intonation features carried by the initial audio segments.

[0078] In one possible scenario, the audio provider's intonation can vary due to factors such as vocal cord function, emotion, tone of voice, and external environment. Therefore, an audio dataset possesses intonation characteristics, which are related to the audio provider; different audio providers exhibit different intonation characteristics. Based on the uniqueness of intonation characteristics, the terminal can refer to these characteristics during the extraction of voiceprint features for audio recognition, thereby improving the accuracy of the extracted voiceprint and ultimately enhancing the accuracy of audio recognition.

[0079] However, the intonation changes in a continuous audio stream are gradual, which can lead to the initial intonation features of the initial audio being indistinct. Therefore, the terminal needs to enhance the initial intonation features carried by the initial audio. One method of enhancement is to concatenate disjoint initial audio segments to obtain a concatenated audio segment carrying the target intonation features. The target intonation features carried in the concatenated audio segment are stronger than those in the initial intonation segments, thus enhancing the initial intonation features.

[0080] In one possible scenario, the terminal needs to identify discontinuous initial audio from multiple initial audio sources. The determination process might be as follows: for any given initial audio, calculate the amplitude at the beginning and end of that initial audio; based on the discontinuity of the amplitude at the beginning and end of any initial audio from other initial audio sources, determine that any given initial audio is discontinuous from other initial audio sources. The amplitude value can be used to represent volume, etc. Since audio features such as volume and speech rate change continuously in a continuous audio stream, the continuity of the initial audio can be determined by calculating the amplitude values ​​used to indicate these audio features.

[0081] Optionally, the terminal selects discontinuous initial audio from multiple initial audio sources for splicing. For example... Figure 3 As shown, after acquiring audio data input through the microphone, the terminal performs audio processing to obtain initial audio. This process is repeated multiple times to obtain multiple initial audio samples, which are then concatenated to obtain a spliced ​​audio sample. The multiple initial audio samples acquired by the terminal may be entirely discontinuous, or some of them may be discontinuous.

[0082] The terminal can select any two initial audio files from multiple initial audio files for concatenation. The two initial audio files can be in non-contiguous time sequences, and the terminal can also select any number of initial audio files for concatenation. Furthermore, the terminal adjusts the concatenation order during the concatenation process. For example, when concatenating initial audio files A and B, the terminal will concatenate the end of initial audio file A with the beginning of initial audio file B, and vice versa, resulting in two concatenated audio files. That is, when the number of initial audio files is a positive integer N greater than 1, and the concatenation method is pairwise concatenation, the number of concatenated audio files is N(N-1). By concatenating multiple initial audio files pairwise, each initial audio file is concatenated with other initial audio files, resulting in at least one concatenated audio file.

[0083] For example, when a terminal splices discontinuous initial audio, it means that the spliced ​​initial audio is not connected at the audio splicing point. For instance, initial audio 1 provides speech data of the subject speaking between 00:00 and 00:10, and initial audio 2 provides speech data of the subject speaking between 00:10 and 00:20. Although initial audio 1 and initial audio 2 are continuous speech data, when the terminal splices the end of initial audio 2 to the beginning of initial audio 1, the splicing points belong to the 20s and 0s of the speech data respectively, and are not connected. Therefore, the splicing of initial audio 1 and initial audio 2 also falls under the category of splicing discontinuous initial audio. By splicing discontinuous initial audio, a jump in intonation is achieved at the audio splicing point, amplifying the change in intonation and enhancing the initial intonation features at the audio splicing point, thus obtaining the target intonation features.

[0084] In step 203, the voiceprint of the reference audio is obtained, and the object information corresponding to the reference audio is known.

[0085] For example, the terminal acquires the voiceprint of the reference audio as the second voiceprint, with one reference audio voiceprint corresponding to one second voiceprint. The process of the terminal acquiring the voiceprint of the reference audio includes, but is not limited to: acquiring the reference audio; performing audio segmentation and voiceprint extraction operations on the reference audio to obtain multiple second voiceprints. The process of the terminal acquiring the reference audio is similar to the process of acquiring the initial audio, and can be referred to the relevant description in step 201, which will not be repeated here. Knowing the object information of the reference audio means that the object information of the object providing the reference audio is in a known state.

[0086] Taking the reference audio obtained through an audio input interface as an example, when the terminal obtains the reference audio, it also provides an information input control. The object providing the reference audio will also input its own information during the input process, thus allowing the terminal to obtain the reference audio with known object information. Alternatively, an audio matching library storing the reference audio may also store the object information corresponding to the reference audio. When the terminal accesses the audio matching library to obtain the reference audio, it will also obtain the corresponding object information.

[0087] The terminal can also directly obtain the second voiceprint. The terminal accesses a verification voiceprint library, which stores multiple pre-processed second voiceprints. Regardless of whether the second voiceprint is obtained in advance or during the audio recognition process, the processing of the second voiceprint is similar to that of the first voiceprint. It involves splicing disjoint reference audio, then segmenting the spliced ​​audio to obtain multiple second audio tracks. The voiceprint of each second audio track is then extracted to obtain the second voiceprint. For example... Figure 4 The process is shown below. For a detailed description, please refer to step 204, which describes the process of obtaining the first voiceprint; it will not be repeated here.

[0088] In step 204, the audio recognition results of multiple initial audios are determined based on the voiceprints of multiple spliced ​​audios and the reference audio.

[0089] Since multiple initial audio files are actually provided by a single audio provider, when the object information corresponding to any one initial audio file is known, the object information of the other initial audio files is also known. Therefore, determining the audio recognition result of multiple initial audio files is equivalent to obtaining the object information of the audio provider through multiple initial audio files.

[0090] For example, the process of determining the audio recognition result includes: acquiring multiple first voiceprints of multiple concatenated audios, where each concatenated audio corresponds to at least one first voiceprint, and at least one first voiceprint in the first voiceprint corresponding to the concatenated audio is determined based on the target intonation features of the audio provider; matching the multiple first voiceprints with multiple second voiceprints to obtain audio matching results of multiple initial audios and a reference audio, where the second voiceprint is the voiceprint of the reference audio, and at least one second voiceprint in the multiple second voiceprints is determined based on the target intonation features corresponding to the reference audio; and determining the audio recognition result of the multiple initial audios based on the audio matching results.

[0091] Since the process of extracting the first voiceprint from multiple concatenated audio files is similar, the following explanation will focus on the process of obtaining the first voiceprint from one concatenated audio file. The voiceprint extraction process for other concatenated audio files can be found in relevant descriptions and will not be repeated here. Optionally, the process of extracting the first voiceprint includes: performing audio segmentation on any concatenated audio file to obtain at least one first audio file; the at least one first audio file contains audio splicing points where the first audio file includes the concatenated audio files, and the audio splicing points carry the target intonation features of the audio provider; and extracting the voiceprints of each of the at least one first audio file to obtain multiple first voiceprints.

[0092] In one possible implementation, the terminal uses a moving window to segment the spliced ​​audio to obtain a first audio with a fixed length. In this case, the audio segmentation process includes: determining the length and step size of the moving window used for audio segmentation; and segmenting any spliced ​​audio based on the length and step size of the moving window, with reference to the audio splicing point of any spliced ​​audio, to obtain at least one first audio.

[0093] Optionally, the length w of the moving window is determined based on the frame length x, for example, w = 2x. The frame length x is a value set empirically. Audio segmentation of spliced ​​audio by referencing any splicing point means taking audio frames of length x to the left and right from the splicing point as the starting point, using these frames as the moving window, and then segmenting the audio. For example, see attached... Figure 5 As shown in (1), Figure 5In the audio format, "A length audio" represents the left audio segment of length A, and "B length audio" represents the audio segment of length B. The left and right audio segments are distinguished by their positions at the audio splicing point. The left audio segment is the one to the left of the splicing point, and the right audio segment is the one to the right. The left and right audio segments can be the same length or different lengths. Figure 5 In (1), the moving window includes a portion of the audio in the left audio and a portion of the audio in the right audio. Furthermore, since the terminal segments the audio based on the audio splicing point, the first audio segment will include the audio splicing point of the left and right audio. Figure 5 The location of the audio splicing points in (1), (2), (4), (5), and (6) can be found in [reference]. Figure 5 The location of the audio splicing point in (3).

[0094] After determining the starting position of the moving window in the audio splicing, the terminal will begin moving the moving window to segment the audio. Optionally, the terminal can move the moving window left and right to segment the left and right audio segments separately. Figure 5 The process shown in (2)-(4).

[0095] Figure 6 This is a schematic diagram of an audio segmentation process provided in an embodiment of this application. The following will illustrate this process. Figure 6 This section provides an example illustrating the process of audio segmentation using a moving window. (See also...) Figure 6 Before segmenting the spliced ​​audio using a moving window, the terminal also calculates the length of the initial audio in the spliced ​​audio, that is... Figure 6 The process of calculating the lengths of the left and right audio segments is described. Optionally, the terminal can calculate the length of the initial audio segments during the acquisition process. For example, if the terminal has a timing function, it can use this function to calculate the length of the audio segments during the process of providing input audio data to the audio provider. The terminal can also use other methods to calculate the length of each initial audio segment in other results.

[0096] Regardless of the method used by the terminal to determine the length of the initial audio, after determining the length of the moving window, it can determine the relationship between the length of the moving window and the length of the initial audio corresponding to the spliced ​​audio. The initial audio corresponding to the spliced ​​audio refers to the initial audio obtained by splicing the audio; for example, the left and right audio in the above embodiment are the initial audio corresponding to the spliced ​​audio. When the length of the moving window is greater than both the length of the left and right audio, audio segmentation using the moving window is not necessary; the left and right audio can be directly taken. Therefore, the terminal can determine that the process of audio segmentation using the moving window has ended. When the moving window is not greater than both the length of the left and right audio (i.e., at least one of the left and right audio has a length greater than the length of the moving window), the terminal begins to use the moving window for left and right segmentation.

[0097] For example, the process of segmenting audio to the left by the terminal includes determining the relationship between the length of the moving window and the length of the left audio. If the length of the moving window is greater than the length of the left audio in the concatenated audio, the process of moving the moving window to the left ends. If the length of the moving window is not greater than the length of the left audio in the concatenated audio, the process of moving the moving window to the left begins to segment the audio.

[0098] Optionally, the moving window process includes: determining the starting point of the moving window, determining the ending point of the moving window based on its length, and cutting the audio data between the starting and ending points to obtain the first audio. If the moving window is shorter than the length of the left audio in the spliced ​​audio, when the terminal starts moving the moving window to the left, the terminal determines the starting point of the moving window as s1 = M + w / 2, that is, taking a position one frame length to the right of M as the starting point of the moving window, where M refers to the audio splicing point. The terminal determines the ending point of the moving window as e1 = s1 - w, that is, taking a position one moving window length to the left of the starting point of the moving window as the ending point of the moving window.

[0099] After determining the endpoint of the moving window, the terminal can check if the endpoint e1 is greater than 0. If e1 is not greater than 0, meaning the moving window has reached the start point of the left audio, the audio data with the start point s1 and the endpoint 0 is segmented to obtain the first audio, and this moving window is used to complete the left audio segmentation. If e1 is greater than 0, meaning the moving window has not reached the start point of the left audio, the audio data with the start point s1 and the endpoint e1 is segmented to obtain the first audio. Furthermore, if the moving window has not reached the start point of the left audio, the terminal can continue to move the moving window to the left, updating the start point s1 = s1 - y according to the step size y of the moving window. The step size is subtracted from the original start point to obtain a new start point, and the above audio segmentation operation using the moving window is repeated based on the new start point. The step size is a value set based on experience, and the step size can be the same as or different from the frame length.

[0100] In addition to moving the moving window to the left to segment the left audio, the terminal also moves the moving window to the right to segment the right audio. The process of segmenting the right audio is similar to that of segmenting the left audio, involving determining the relationship between the length of the moving window and the length of the right audio, and then deciding whether to perform audio segmentation based on this relationship. The difference lies in the fact that during the right audio segmentation, the starting point of the moving window is s2 = Mw / 2, meaning the audio splicing point is moved to the left by one frame length, and the ending point of the moving window is e2 = s2 + w, meaning the starting point of the moving window is moved to the right by one moving window length. The right audio segmentation process involves updating the starting point s2 = s2 + y based on the step size y of the moving window. For a more detailed description, please refer to the relevant description of left audio segmentation in the above embodiment, which will not be repeated here.

[0101] By using a fixed-length moving window to segment the concatenated audio, multiple first audio segments of uniform length are obtained, thus transforming the variable-length audio recognition task into a fixed-length audio recognition task and reducing the difficulty of audio recognition. Since the audio segmentation process using the moving window references the audio splicing point, at least one of the multiple first audio segments will include the audio splicing point. For example, the first audio segment obtained from the initial segmentation using the moving window will include the audio splicing point because the step size y is taken x to the left and right from the audio segmentation point. Subsequent segmentations, such as when the moving window moves its step size y is less than x, will result in a second or third segmentation, also including the audio splicing point.

[0102] For example, after the terminal completes audio segmentation using a moving window, it adjusts the length of the moving window according to a length increment to obtain a longer first audio segment. The length increment is a value set based on experience. For instance, if the length increment is set to 1 based on experience, the length of the next moving window will be w = 2(x + 1), that is... Figure 5 The situation shown in (5) is as follows. By adjusting the length of the moving window, audio segmentation is performed based on the new moving window. The condition for ending the above audio segmentation using the moving window is that the length of the moving window is not less than the length of the spliced ​​audio, that is... Figure 5 The situation in (6).

[0103] Optionally, the number of moving windows can be one or more. When there are multiple moving windows used for audio segmentation, the length of any one of the multiple moving windows is different from that of the other moving windows. The other moving windows are the moving windows other than any one of the multiple moving windows.

[0104] In one possible scenario, after the terminal completes audio segmentation using a moving window, it will perform a filtering operation on the segmented first audio, filtering out duplicate first audio in the first audio, so that each first audio in the multiple filtered first audios is different.

[0105] After completing audio segmentation to obtain multiple first audio samples, the terminal can extract the voiceprint of each first audio sample to obtain the first voiceprint, for example... Figure 7 As shown. The embodiments of this application do not limit the method for extracting the voiceprint of the first audio. The voiceprint can be extracted using the MFCC (Mel Frequency Cepstrum Coefficient) algorithm, or other voiceprint extraction algorithms.

[0106] For example, after acquiring the first voiceprint, the terminal can calculate the similarity between each of the multiple first voiceprints and each of the multiple second voiceprints; based on the similarity between each first voiceprint and each second voiceprint and a matching threshold, the audio matching result is determined. This application does not limit the process by which the terminal calculates the similarity between the first and second voiceprints; the terminal can use cosine similarity or PLDA (Probabilistic Linear Discriminant Analysis) to calculate the similarity between the first and second voiceprints.

[0107] After calculating the similarity between the first and second voiceprints, the terminal can compare this similarity with a matching threshold to determine the audio matching result. For example, if the similarity between any first voiceprint and any second voiceprint is greater than the matching threshold, the audio matching result is determined to be a successful match; or, if the similarity between any first voiceprint and any second voiceprint is not greater than the matching threshold, the audio matching result is determined to be a failed match. The matching threshold can be set based on experience, for example, a matching threshold of 90%.

[0108] Since the similarity between one first voiceprint and one second voiceprint is greater than the matching threshold, the terminal can determine that the initial audio and the reference audio have matched successfully, and there is no need to continue matching other first voiceprints and other second voiceprints. Therefore, the terminal can choose to first calculate the similarity between each first voiceprint and second voiceprint, and then obtain the relationship between the similarity and the matching threshold. Alternatively, it can... Figure 8 As shown, similarity calculations and comparisons are performed sequentially. For example, first, a first voiceprint and a second voiceprint are selected. The similarity between the first and second voiceprints is calculated. Then, the similarity between the first and second voiceprints is compared with a matching threshold. If the similarity is not greater than the matching threshold, it is determined that the first and second voiceprints have failed to match. A new second voiceprint is selected, and the similarity between the first and second voiceprints is calculated again. After the first voiceprint fails to match with all second voiceprints, a new first voiceprint is selected, and the above operation is repeated. This continues until a first voiceprint and a second voiceprint have a similarity greater than the matching threshold. At this point, the terminal determines that the match is successful and the matching operation between the first and second voiceprints ends.

[0109] Regardless of the order in which the terminal calculates the similarity between the first and second voiceprints and compares them with the matching threshold, the audio recognition result can be determined based on the audio matching result. If the audio matching result is successful, the object information corresponding to the reference audio is determined as the object information of the audio providing object, resulting in audio recognition results for multiple initial audio samples.

[0110] The successful match between the initial audio and the reference audio indicates that they belong to the same audio provider. Since the object information corresponding to the reference audio is known, this object information can be identified as the object information of the audio provider for the initial audio. The terminal thus obtains the object information of the audio provider for the initial audio, thereby completing audio recognition.

[0111] In summary, the audio recognition method provided in this application, by splicing disjoint initial audio, causes a pitch shift at the splicing point, resulting in a target pitch feature with stronger characteristics than the initial pitch feature. By using the stronger target pitch feature to assist audio recognition, the accuracy of audio recognition is improved, effectively addressing the problem of inaccurate voiceprint recognition caused by the phonemes of the audio source.

[0112] Furthermore, in the process of audio recognition using audio matching, multiple first voiceprints are matched with multiple second voiceprints. By increasing the number of matching operations, the accuracy of voiceprint matching is improved, thereby increasing the accuracy of audio recognition results. Since audio segmentation is performed based on a moving window, it ensures that multiple first audio segments obtained from the same moving window have the same length, thus transforming the recognition of variable-length audio into the recognition of fixed-length audio, reducing the difficulty and cost of audio recognition.

[0113] See Figure 9 This application provides an audio recognition device, which includes:

[0114] The acquisition module 901 is used to acquire multiple initial audios, which correspond to the same audio provider object. The object information of the audio provider object is unknown.

[0115] The splicing module 902 is used to splice unconnected initial audio from multiple initial audios to obtain multiple spliced ​​audios. The target intonation features carried by the spliced ​​audio are stronger than the initial intonation features carried by the initial audio.

[0116] The acquisition module 901 is also used to acquire the voiceprint of the reference audio, and the object information corresponding to the reference audio is known;

[0117] The determination module 903 is used to determine the audio recognition results of multiple initial audios based on the voiceprints of multiple spliced ​​audios and reference audios.

[0118] Optionally, the determining module 903 is used to acquire multiple first voiceprints of multiple concatenated audios, where each concatenated audio corresponds to at least one first voiceprint, and at least one first voiceprint among the first voiceprints corresponding to a concatenated audio is determined based on the target intonation features of the audio provider; the multiple first voiceprints are matched with multiple second voiceprints to obtain audio matching results of multiple initial audios and a reference audio, where the second voiceprint is the voiceprint of the reference audio, and at least one second voiceprint among the multiple second voiceprints is determined based on the target intonation features corresponding to the reference audio; and the audio recognition results of the multiple initial audios are determined based on the audio matching results.

[0119] Optionally, the determining module 903 is used to perform audio segmentation on any spliced ​​audio to obtain at least one first audio, wherein the at least one first audio contains an audio splicing point of the spliced ​​audio, and the audio splicing point carries the target intonation features of the audio providing object; and to extract the voiceprint of each of the at least one first audio to obtain a first voiceprint.

[0120] Optionally, the determining module 903 is used to determine the length and step size of the moving window for audio segmentation; based on the length and step size of the moving window, the audio segmentation is performed on any spliced ​​audio with reference to the audio splicing point of any spliced ​​audio to obtain at least one first audio.

[0121] Optionally, there are multiple moving windows used for audio segmentation. The length of any one of the multiple moving windows is different from the length of the other moving windows. The other moving windows are the moving windows other than any one of the multiple moving windows.

[0122] Optionally, the determining module 903 is used to calculate the similarity between each of the first voiceprints in the plurality of first voiceprints and each of the second voiceprints in the plurality of second voiceprints; and to determine the audio matching result based on the similarity between each of the first voiceprints and each of the second voiceprints and the matching threshold.

[0123] Optionally, the determining module 903 is used to determine the audio matching result as a successful match based on the fact that the similarity between any first voiceprint and any second voiceprint is greater than the matching threshold; or, to determine the audio matching result as a failed match based on the fact that the similarity between each first voiceprint and each second voiceprint is not greater than the matching threshold.

[0124] Optionally, the determining module 903 is used to determine the object information corresponding to the reference audio as the object information of the audio providing object based on the successful audio matching result, thereby obtaining the audio recognition results of multiple initial audios.

[0125] The aforementioned device splices disjoint initial audio segments, creating a dramatic shift in intonation in the spliced ​​audio. This spliced ​​audio carries stronger target intonation features than the initial intonation features of the unspliced ​​audio. This stronger target intonation feature assists in audio recognition, improving the accuracy of audio recognition.

[0126] It should be noted that the apparatus provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.

[0127] Figure 10 This is a schematic diagram of a server structure provided in an embodiment of this application. The server can vary significantly due to differences in configuration or performance. It may include one or more processors 1001 and one or more memories 1002. The one or more memories 1002 store at least one computer program, which is loaded and executed by the one or more processors 1001 to enable the server to implement the audio recognition methods provided in the various method embodiments described above. The processor 1001 is, for example, a Central Processing Unit (CPU). Of course, the server may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server may also include other components for implementing device functions, which will not be elaborated here.

[0128] Figure 11 This is a schematic diagram of the structure of an audio recognition device provided in an embodiment of this application. The device can be a terminal, such as a smartphone, tablet computer, MP3 (Moving Picture Experts Group Audio Layer III) player, MP4 (Moving Picture Experts Group Audio Layer IV) player, laptop computer, or desktop computer. The terminal may also be referred to as a user device, portable terminal, laptop terminal, desktop terminal, or other names.

[0129] Typically, a terminal includes a processor 1101 and a memory 1102.

[0130] Processor 1101 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1101 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1101 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1101 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 1101 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0131] The memory 1102 may include one or more computer-readable storage media, which may be non-transitory. The memory 1102 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1102 are used to store at least one instruction, which is executed by the processor 1101 to cause the terminal to implement the audio recognition method provided in the method embodiments of this application.

[0132] In some embodiments, the terminal may also optionally include: a peripheral device interface 1103 and at least one peripheral device. The processor 1101, memory 1102, and peripheral device interface 1103 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1103 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of: a radio frequency circuit 1104, a display screen 1105, a camera assembly 1106, an audio circuit 1107, a positioning assembly 1108, and a power supply 1109.

[0133] Peripheral device interface 1103 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 1101 and memory 1102. In some embodiments, processor 1101, memory 1102 and peripheral device interface 1103 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 1101, memory 1102 and peripheral device interface 1103 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0134] The radio frequency (RF) circuit 1104 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1104 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1104 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 1104 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 1104 can communicate with other terminals through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: metropolitan area networks (MANs), various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks (WLANs), and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1104 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.

[0135] Display screen 1105 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 1105 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 1101 for processing. In this case, display screen 1105 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, display screen 1105 can be a single screen, located on the front panel of the terminal; in other embodiments, display screen 1105 can be at least two screens, respectively located on different surfaces of the terminal or in a folded design; in other embodiments, display screen 1105 can be a flexible display screen, located on a curved or folded surface of the terminal. Furthermore, display screen 1105 can be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. Display screen 1105 can be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).

[0136] The camera assembly 1106 is used to acquire images or videos. Optionally, the camera assembly 1106 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 1106 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.

[0137] The audio circuit 1107 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 1101 for processing, or input to the radio frequency circuit 1104 to achieve voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each positioned at a different location on the terminal. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 1101 or the radio frequency circuit 1104 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 1107 may also include a headphone jack.

[0138] The positioning component 1108 is used to locate the current geographical location of the terminal to enable navigation or LBS (Location Based Service). The positioning component 1108 can be a positioning component based on the US GPS (Global Positioning System), China's BeiDou system, Russia's GLONASS system, or the EU's Galileo system.

[0139] Power supply 1109 is used to power the various components in the terminal. Power supply 1109 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 1109 includes a rechargeable battery, the rechargeable battery can support wired charging or wireless charging. The rechargeable battery can also be used to support fast charging technology.

[0140] In some embodiments, the terminal further includes one or more sensors 1110. The one or more sensors 1110 include, but are not limited to: an accelerometer 1111, a gyroscope 1112, a pressure sensor 1113, a fingerprint sensor 1114, an optical sensor 1115, and a proximity sensor 1116.

[0141] Accelerometer 1111 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by the terminal. For example, accelerometer 1111 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 1101 can control display screen 1105 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 1111. Accelerometer 1111 can also be used for collecting game or user motion data.

[0142] The gyroscope sensor 1112 can detect the terminal's orientation and rotation angle. The gyroscope sensor 1112 can work in conjunction with the accelerometer sensor 1111 to collect the user's 3D movements on the terminal. Based on the data collected by the gyroscope sensor 1112, the processor 1101 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.

[0143] The pressure sensor 1113 can be disposed on the side bezel of the terminal and / or the lower layer of the display screen 1105. When the pressure sensor 1113 is disposed on the side bezel of the terminal, it can detect the user's grip signal on the terminal, and the processor 1101 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 1113. When the pressure sensor 1113 is disposed on the lower layer of the display screen 1105, the processor 1101 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 1105. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0144] The fingerprint sensor 1114 is used to collect the user's fingerprint. The processor 1101 identifies the user's identity based on the fingerprint collected by the fingerprint sensor 1114, or the fingerprint sensor 1114 identifies the user's identity based on the collected fingerprint. When the user's identity is identified as trusted, the processor 1101 authorizes the user to perform relevant sensitive operations, including unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings. The fingerprint sensor 1114 can be located on the front, back, or side of the terminal. When the terminal has physical buttons or a manufacturer's logo, the fingerprint sensor 1114 can be integrated with the physical buttons or manufacturer's logo.

[0145] An optical sensor 1115 is used to collect ambient light intensity. In one embodiment, the processor 1101 can control the display brightness of the display screen 1105 based on the ambient light intensity collected by the optical sensor 1115. Specifically, when the ambient light intensity is high, the display brightness of the display screen 1105 is increased; when the ambient light intensity is low, the display brightness of the display screen 1105 is decreased. In another embodiment, the processor 1101 can also dynamically adjust the shooting parameters of the camera assembly 1106 based on the ambient light intensity collected by the optical sensor 1115.

[0146] The proximity sensor 1116, also known as a distance sensor, is typically installed on the front panel of the terminal. The proximity sensor 1116 is used to detect the distance between the user and the front of the terminal. In one embodiment, when the proximity sensor 1116 detects that the distance between the user and the front of the terminal is gradually decreasing, the processor 1101 controls the display screen 1105 to switch from a screen-on state to a screen-off state; when the proximity sensor 1116 detects that the distance between the user and the front of the terminal is gradually increasing, the processor 1101 controls the display screen 1105 to switch from a screen-off state to a screen-on state.

[0147] Those skilled in the art will understand that Figure 11 The structure shown does not constitute a limitation on the audio recognition device and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0148] In an exemplary embodiment, a computer device is also provided, comprising a processor and a memory storing at least one computer program. The at least one computer program is loaded and executed by one or more processors to enable the computer device to implement any of the above-described audio recognition methods.

[0149] In an exemplary embodiment, a computer-readable storage medium is also provided, which stores at least one computer program that is loaded and executed by a processor of a computer device to enable the computer to implement any of the above-described audio recognition methods.

[0150] In one possible implementation, the aforementioned computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.

[0151] In an exemplary embodiment, a computer program product or computer program is also provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform any of the above-described audio recognition methods.

[0152] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the initial audio involved in this application was obtained with full authorization.

[0153] It should be understood that "multiple" as used in this article refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0154] The above description is merely an exemplary embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.

Claims

1. An audio recognition method, characterized in that, The method includes: Multiple initial audio files are obtained, each corresponding to the same audio provider object, and the object information of the audio provider object is unknown. By splicing together the disjoint initial audios from the multiple initial audios, multiple spliced ​​audios are obtained, wherein the target intonation features carried by the spliced ​​audios are stronger than the initial intonation features carried by the initial audios. The voiceprint of the reference audio is obtained, and the object information corresponding to the reference audio is known; For any concatenated audio, the audio is segmented starting from the audio concatenation point to obtain at least one first audio. The at least one first audio contains an audio concatenation point that includes the audio of any concatenated audio. The audio concatenation point carries the target intonation features of the audio providing object. Extract the voiceprints of each of the at least one first audio audios to obtain at least one first voiceprint corresponding to any spliced ​​audio audio. At least one first voiceprint in the first voiceprint corresponding to any spliced ​​audio audio is determined according to the target intonation features of the audio provider. The first voiceprint of multiple spliced ​​audios is matched with multiple second voiceprints to obtain the audio matching result of the multiple initial audios and the reference audio. The second voiceprint is the voiceprint of the reference audio. At least one of the multiple second voiceprints is determined according to the target intonation feature corresponding to the reference audio. The audio recognition results of the multiple initial audios are determined based on the audio matching results.

2. The method according to claim 1, characterized in that, The step of segmenting any spliced ​​audio file using the audio splicing point as the starting point to obtain at least one first audio file includes: Determine the length of the moving window used for audio segmentation and the step size of the moving window; Based on the length and step size of the moving window, the audio of any spliced ​​audio is segmented starting from the audio splicing point of any spliced ​​audio to obtain at least one first audio.

3. The method according to claim 2, characterized in that, The moving windows used for audio segmentation are multiple, and the length of any one of the multiple moving windows is different from the length of the other moving windows. The other moving windows are the moving windows other than any one of the multiple moving windows.

4. The method according to any one of claims 1-3, characterized in that, The step of matching the first voiceprint of multiple spliced ​​audio files with multiple second voiceprints to obtain the audio matching results of the multiple initial audio files and the reference audio file includes: Calculate the similarity between each first voiceprint in the first voiceprint of the plurality of spliced ​​audio and each second voiceprint in the plurality of second voiceprints; The audio matching result is determined based on the similarity and matching threshold between each of the first voiceprints and each of the second voiceprints.

5. The method according to claim 4, characterized in that, The step of determining the audio matching result based on the similarity and matching threshold between each first voiceprint and each second voiceprint includes: Based on the fact that the similarity between any first voiceprint and any second voiceprint is greater than the matching threshold, the audio matching result is determined to be a successful match. Alternatively, based on the fact that the similarity between each of the first voiceprints and each of the second voiceprints is not greater than the matching threshold, the audio matching result is determined to be a matching failure.

6. The method according to any one of claims 1-3, characterized in that, The step of determining the audio recognition result of the plurality of initial audios based on the audio matching result includes: Based on the successful audio matching result, the object information corresponding to the reference audio is determined as the object information of the audio providing object, thus obtaining the audio recognition result of the multiple initial audios.

7. An audio recognition device, characterized in that, The device includes: The acquisition module is used to acquire multiple initial audio files, which correspond to the same audio provider object, and the object information of the audio provider object is unknown. The splicing module is used to splice unconnected initial audios from the plurality of initial audios to obtain a plurality of spliced ​​audios, wherein the target intonation features carried by the spliced ​​audios are stronger than the initial intonation features carried by the initial audios. The acquisition module is also used to acquire the voiceprint of the reference audio, wherein the object information corresponding to the reference audio is known; A determining module is configured to: segment any concatenated audio file starting from an audio splicing point to obtain at least one first audio file, wherein the at least one first audio file contains an audio splicing point of the concatenated audio file, and the audio splicing point carries the target intonation features of the audio provider; extract the voiceprints of each of the at least one first audio file to obtain at least one first voiceprint corresponding to the concatenated audio file, wherein at least one first voiceprint corresponding to the concatenated audio file is determined based on the target intonation features of the audio provider; match the first voiceprints of multiple concatenated audio files with multiple second voiceprints to obtain an audio matching result between the multiple initial audio files and the reference audio file, wherein the second voiceprint is the voiceprint of the reference audio file, and at least one second voiceprint is determined based on the target intonation features corresponding to the reference audio file; and determine the audio recognition result of the multiple initial audio files based on the audio matching result.

8. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one computer program, which is loaded and executed by the processor to enable the computer device to implement the audio recognition method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Voiceprint recognition method and device, computer equipment and storage medium

    CN112820297A

  • Audio recognition method and device, equipment and storage medium

    CN114386006A

  • Voiceprint recognition model training method and device and storage medium

    CN114420136A